AI 大语言模型目录
青柠AI 现已汇聚 324 款前沿语言与多模态模型,统一提供 OpenAI 格式接口,支持 Cursor、Claude Code 等开发工具免外卡直连。
模型目录与 API 接入参数事实清单 (Models Fact Sheet)
已收录 324 款模型 · 涵盖全部前沿大语言模型与多模态模型
Alibaba 系列
53 个模型Qwen vision-language model for visual reasoning, documents, and agent tasks
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Flagship Qwen model for complex reasoning, coding, and agentic workflows
Translation model for multilingual conversion, localization, and cross-language workflows
Translation model for multilingual conversion, localization, and cross-language workflows
Qwen omni model for text, vision, audio, and multimodal agent tasks
Qwen omni model for text, vision, audio, and multimodal agent tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen instruction model for multilingual chat, reasoning, and tool use
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Qwen vision-language model for visual reasoning, documents, and agent tasks
OCR model for extracting structured text from documents and screenshots
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen omni model for text, vision, audio, and multimodal agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen instruction model for multilingual chat, reasoning, and tool use
Large open Qwen MoE for multilingual reasoning, coding, and tool use
Dense open Qwen model for self-hosted chat, reasoning, and coding
Qwen instruction model for multilingual chat, reasoning, and tool use
Speech transcription model for accurate audio-to-text and captioning workflows
Smaller Qwen coder for efficient local agents and repo-level fixes
Open Qwen coding heavyweight for repository reasoning and agentic engineering
Qwen coding model for software agents, repository edits, and code reasoning
Hosted Qwen coder for software agents, repo edits, and long-context code
Speech generation model for controllable voice, narration, and audio delivery
Flagship Qwen3 model for coding agents, complex reasoning, and tool use
Qwen instruction model for multilingual chat, reasoning, and tool use
Efficient Qwen thinking model for local reasoning, math, and coding agents
Qwen omni model for text, vision, audio, and multimodal agent tasks
Qwen omni model for text, vision, audio, and multimodal agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Large open Qwen multimodal MoE for visual agents and long technical tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Qwen vision-language model for visual reasoning, documents, and agent tasks
Open multimodal Qwen MoE for local agents that need vision, audio, and code
Qwen vision-language model for visual reasoning, documents, and agent tasks
Flagship Qwen model for complex reasoning, coding, and agentic workflows
Earlier Qwen multimodal workhorse for million-token agent and document tasks
Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks
Multimodal Qwen workhorse for long-context agents, visual inputs, and coding
Qwen vision-language model for visual reasoning, documents, and agent tasks
2.4-trillion-parameter MoE flagship for coding, professional work, multimodal understanding, and long-horizon agentic workflows
Qwen reasoning model for deliberate problem solving, math, and coding
AlibabaCn 系列
18 个模型Qwen instruction model for multilingual chat, reasoning, and tool use
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen instruction model for multilingual chat, reasoning, and tool use
Efficient Qwen model for fast chat, extraction, and high-volume workloads
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen coding model for software agents, repository edits, and code reasoning
Qwen coding model for software agents, repository edits, and code reasoning
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen instruction model for multilingual chat, reasoning, and tool use
Qwen vision-language model for visual reasoning, documents, and agent tasks
Lightweight multimodal Qwen model for high-throughput text, image, and video tasks
Qwen reasoning model for deliberate problem solving, math, and coding
DeepSeek reasoning model for multi-step analysis, math, coding, and tools
DeepSeek chat model for instruction following, coding, and analysis
DeepSeek chat model for instruction following, coding, and analysis
DeepSeek chat model for instruction following, coding, and analysis
General-purpose chat model for instruction following, writing, and analysis
AlibabaCodingPlan 系列
2 个模型AlibabaTokenPlan 系列
8 个模型Video model for image-to-video generation
Video model for reference-guided video generation
Video model for prompt-driven text-to-video generation
Image model for prompt-driven generation, editing, and visual design workflows
Image model for prompt-driven generation, editing, and visual design workflows
Preview Qwen flagship for million-token multimodal reasoning and long-horizon agentic workflows
Image model for prompt-driven generation, editing, and visual design workflows
Image model for prompt-driven generation, editing, and visual design workflows
Anthropic 系列
14 个模型Claude model for creative writing, analysis, and controlled agent workflows
Claude model for demanding reasoning and long-horizon agentic work
Fast Claude lane for lightweight agents, office tasks, and responsive chat
Fast Claude model for responsive assistance, classification, and lightweight agents
Flagship Claude model for deep reasoning, coding, and long-horizon agents
Flagship Claude model for deep reasoning, coding, and long-horizon agents
High-end Claude for difficult coding, planning, and slower expert reasoning
Stronger Opus tier for advanced software work and high-stakes reasoning
Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents
Strongest Claude Opus model for coding, agents, and professional work
Balanced Claude model for coding, analysis, agent workflows, and cost control
Balanced Claude model for coding, analysis, agent workflows, and cost control
Claude workhorse for coding agents, careful analysis, and production cost control
Everyday Claude agent model for coding, planning, browsing, and general work
Bailing 系列
2 个模型Cohere 系列
10 个模型Cohere command model for multilingual enterprise agents, tools, and chat
Cohere's stronger command model for multilingual agents and enterprise workflows
Cohere reasoning model for multilingual enterprise agents, tools, and complex workflows
Translation model for multilingual conversion, localization, and cross-language workflows
Cohere vision model for multilingual document analysis, OCR, and image understanding
Cohere retrieval model for long-context chat and enterprise RAG workflows
Cohere's RAG workhorse for long-context enterprise search and tool use
Cohere retrieval model for long-context chat and enterprise RAG workflows
Open Command R model optimized for Arabic enterprise chat, RAG, and cultural knowledge
Cohere coding model for practical software engineering and agentic edits
DeepSeek 系列
4 个模型DeepSeek V4.1 Flash model for reasoning and agentic coding
DeepSeek V4.1 Flash model for reasoning and agentic coding
DeepSeek V4.1 Flash model for reasoning and agentic coding
DeepSeek V4 Pro snapshot with million-token context and support for thinking and non-thinking modes
Google 系列
34 个模型Maximum-comprehensiveness agentic researcher for multi-step investigation, synthesis, and cited reports
Agentic model for autonomous multi-step research, synthesis, and cited reports
Specialized Gemini 2.5 model for browser-control agents that automate UI tasks
Fast Gemini workhorse for multimodal apps where latency and price matter
Nano Banana image model for fast generation, edits, and character-consistent assets
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
Speech generation model for controllable voice, narration, and audio delivery
Google's proven reasoning model for coding, math, and multimodal analysis
Speech generation model for controllable voice, narration, and audio delivery
New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs
Nano Banana Pro for higher-fidelity image generation and design-heavy edits
Nano Banana Pro for higher-fidelity image generation and design-heavy edits
Image model for prompt-driven generation, editing, and visual design workflows
Image model for prompt-driven generation, editing, and visual design workflows
Low-latency Gemini model for high-volume multimodal and agent workloads
Fastest, most cost-efficient Gemini image model for high-volume 1K generation and editing
Legacy model retained for compatibility with older integrations
High-quality, low-latency Live API model for real-time dialogue and voice-first AI applications
Low-latency speech generation with steerable prompts and expressive audio tags
Reasoning-first Gemini preview for agentic coding and complex problem solving
Advanced Gemini model for complex reasoning, coding, and multimodal analysis
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Low-latency audio-to-audio model for real-time speech translation across 70+ languages
Fast Gemini model balancing multimodal reasoning, tool use, and cost
High-efficiency Gemini model for agentic workflows, coding, and multimodal reasoning
Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
Multimodal embedding model mapping text, images, video, audio, and PDFs into a unified embedding space
High-efficiency Gemini model for agentic workflows, coding, and multimodal reasoning
Fast Gemini model balancing multimodal reasoning, tool use, and cost
Video generation and editing model for fast, conversational text- and image-to-video workflows
Music generation model for short 30-second clips, loops, and previews from text or image prompts
Music generation model for full-length songs from text or images with vocals and structure
Inception 系列
3 个模型Reasoning model for deliberate analysis, multi-step problem solving, and tool use
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception
Reasoning model for deliberate analysis, multi-step problem solving, and tool use
KimiForCoding 系列
4 个模型Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work
256K-context version of Kimi K3, reducing token consumption for shorter coding sessions
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
Lower-latency Kimi Code variant for interactive edits and coding-agent loops
Llama 系列
7 个模型Open multimodal Llama model for strong reasoning and fast responses
Open multimodal Llama model for long-context analysis and efficient agents
Open multimodal Llama model for strong reasoning and fast responses
Open Llama instruction model for multilingual chat, reasoning, and coding
Open Llama instruction model for multilingual chat, reasoning, and coding
Open multimodal Llama model for strong reasoning and fast responses
Open multimodal Llama model for long-context analysis and efficient agents
Longcat 系列
1 个模型Meta 系列
5 个模型Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.
Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1 with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows.
Muse Spark 1.2 is a coding-focused update to Muse Spark 1.1 with improvements in code generation, complex debugging, codebase understanding, and end-to-end developer workflows.
Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It improves long-horizon agent collaboration, instruction following, and coding efficiency relative to Muse Spark 1.2.
Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It improves long-horizon agent collaboration, instruction following, and coding efficiency relative to Muse Spark 1.2.
Minimax 系列
7 个模型Efficient open MiniMax model built for coding agents and tool-heavy workflows
Earlier MiniMax agent model for practical coding and productivity tasks
Prior MiniMax coding model for agent workflows, office edits, and automation
High-speed MiniMax model for low-latency coding and agent workflows
Open MiniMax flagship for coding agents, office automation, and complex environments
Low-latency M2.7 variant for interactive coding plans and agent loops
MiniMax multimodal model for long-context coding, perception, and agent planning
Mistral 系列
32 个模型Mistral code model for completions, refactors, and developer IDE workflows
Mistral's coding-agent model for repository work, terminal tasks, and software fixes
Legacy model retained for compatibility with older integrations
Legacy model retained for compatibility with older integrations
Legacy model retained for compatibility with older integrations
Legacy model retained for compatibility with older integrations
Legacy model retained for compatibility with older integrations
Legacy model retained for compatibility with older integrations
Mistral reasoning model for transparent analysis, math, and complex decisions
Mistral reasoning model for transparent analysis, math, and complex decisions
Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads
Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Mistral's largest general model for enterprise agents, coding, and multilingual reasoning
Flagship Mistral model for advanced reasoning, coding, and multilingual work
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
Balanced Mistral model for enterprise assistants, multilingual work, and tools
Balanced Mistral model for enterprise assistants, multilingual work, and tools
Efficient Mistral-NVIDIA open model for multilingual chat and local deployment
Efficient Mistral model for fast chat, extraction, and production assistants
Fast Mistral production model for chat, extraction, and cost-sensitive agents
Efficient Mistral model for fast chat, extraction, and production assistants
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
Legacy model retained for compatibility with older integrations
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
Mistral vision-language model for image understanding and multimodal chat
Mistral's larger vision model for document-heavy image understanding and chat
Instruct model with native audio input for speech understanding and tool use
Open flagship GLM for long-horizon coding agents and million-token context work
Moonshotai 系列
4 个模型Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking
Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work
Morph 系列
3 个模型Automatic model router for matching prompts to suitable backends and budgets
Efficient model for low-latency assistance, extraction, and routine automation
Flagship model for demanding analysis, coding, and production agent workflows
Nova 系列
2 个模型OpenAI 系列
44 个模型Compact GPT model for low-latency assistance and high-volume workloads
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Compact GPT model for low-latency assistance and high-volume workloads
Long-lived GPT workhorse for coding, instruction following, and production apps
Affordable GPT-4.1 lane for fast coding help and structured extraction
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
Omni-era GPT for multimodal chat, practical coding, and general assistants
GPT model for general reasoning, writing, coding, and tool-assisted tasks
GPT model for general reasoning, writing, coding, and tool-assisted tasks
GPT model for general reasoning, writing, coding, and tool-assisted tasks
Small omni GPT for cheap multimodal assistance and production-scale traffic
Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows
Small GPT-5 for responsive agents, coding help, and everyday automation
Tiny GPT-5 lane for routing, extraction, classification, and bulk jobs
Higher-accuracy GPT-5 tier for tough analysis, coding reviews, and planning
Sharper GPT-5 generation for coding, product work, and tool-assisted tasks
Reliable GPT generation for broad coding, writing, and tool-assisted product work
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
Higher-accuracy GPT-5.2 variant for tougher reasoning and review workflows
Chat-tuned GPT model for conversational assistance, writing, and tool workflows
Coding-optimized GPT model for repository edits, reviews, and agentic software work
Coding-optimized GPT model for repository edits, reviews, and agentic software work
Agent-ready GPT for coding and computer-use workflows at a lower cost
Strong small GPT for coding subagents, quick tool use, and high-volume work
Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation
More exact GPT-5.4 tier for demanding professional reasoning and agent tasks
Default frontier GPT for coding, computer use, research, and knowledge work
Highest-accuracy GPT-5.5 tier for slower, precision-heavy reasoning and coding
Frontier GPT-5.6 model for complex professional work, coding, and agentic workflows
Cost-efficient GPT-5.6 model for fast, high-volume workloads
Frontier GPT-5.6 model for complex professional work, coding, and agentic workflows
Balanced GPT-5.6 model for capable, cost-efficient everyday work
GPT-6 Astra is OpenAI's most capable model for complex reasoning, coding, computer use, research, and document creation.
Image model for prompt-driven generation, editing, and visual design workflows
Realtime speech-to-speech model with configurable reasoning, tool use, and robust voice-agent behavior
O-series reasoning model for hard analysis, math, coding, and planning
O-series reasoning model for hard analysis, math, coding, and planning
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
Smaller o-series reasoner for economical coding, math, and planning tasks
High-effort o3 tier for difficult technical reasoning and careful answers
Fast o-series model for compact reasoning, coding, and tool use
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
Perplexity 系列
4 个模型Fast web-grounded Sonar for current answers, citations, and lightweight retrieval
Sonar search model for current answers, retrieval, and citation-backed chat
Deeper Sonar search model with broader retrieval and stronger synthesis
Web-grounded Sonar for multi-step research questions that need cited reasoning
Poolside 系列
3 个模型Poolside's open-weight model for agentic coding and long-horizon work
Agentic coding model from Poolside in the XS size class for local deployment
Agentic coding model from Poolside in the XS size class for local deployment
Sakana 系列
3 个模型Quality-first multi-agent model for hard research, analysis, and competitions
Quality-first multi-agent model for hard research, analysis, and competitions
Japanese-specialized reasoning model based on Kimi K2.6 and tuned for Japanese language, culture, and business workflows
Stepfun 系列
5 个模型StepFun flash model for efficient multimodal reasoning, coding, and tool use
StepFun flash model for efficient multimodal reasoning, coding, and tool use
StepFun flash lane for quick multimodal reasoning and coding assistance
StepFun flash model for efficient multimodal reasoning, coding, and tool use
Newer StepFun flash model for faster agents, coding, and multimodal prompts
TencentCodingPlan 系列
5 个模型Tencent Hy reasoning model for coding, instruction following, and agent tasks
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Tencent Hy reasoning model for coding, instruction following, and agent tasks
Automatic model router for matching prompts to suitable backends and budgets
TencentTokenPlan 系列
2 个模型TencentTokenhub 系列
1 个模型Thinkingmachines 系列
2 个模型Upstage 系列
4 个模型Efficient model for low-latency assistance, extraction, and routine automation
Flagship model for demanding analysis, coding, and production agent workflows
Flagship model for demanding analysis, coding, and production agent workflows
Upstage's flagship model, specialized for agentic use
V0 系列
3 个模型Multimodal reasoning model for visual analysis, planning, and tool use
Multimodal reasoning model for visual analysis, planning, and tool use
Multimodal reasoning model for visual analysis, planning, and tool use
Xiaomi 系列
6 个模型Legacy model retained for compatibility with older integrations
Legacy model retained for compatibility with older integrations
Earlier MiMo Pro model for multimodal agents, reasoning, and code tasks
Open MiMo model for multimodal coding agents and long-context automation
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
MiMo pro model for strong multimodal reasoning and agent execution
XiaomiTokenPlanAms 系列
4 个模型Speech generation model for controllable voice, narration, and audio delivery
Speech generation model for controllable voice, narration, and audio delivery
Speech generation model for controllable voice, narration, and audio delivery
Speech generation model for controllable voice, narration, and audio delivery
Zai 系列
16 个模型Hybrid-reasoning GLM release that made the 4.5 line broadly useful
Lighter GLM-4.5 variant for fast coding assistance and cheaper agents
Efficient GLM model for fast reasoning, coding, and agent workflows
GLM vision model for visual reasoning, documents, and multimodal agents
Late GLM-4 workhorse for coding agents, reasoning, and structured tasks
GLM vision model for visual reasoning, documents, and multimodal agents
Mature GLM model for dependable coding, reasoning, and structured agent tasks
Budget GLM lane for fast coding help, routing, and everyday automation
Efficient GLM model for fast reasoning, coding, and agent workflows
General GLM flagship for coding, analysis, and tool-heavy engineering workflows
Faster GLM-5 lane for coding agents that need lower latency
Strong GLM coding model for agentic engineering, terminals, and repository generation
Open flagship GLM for long-horizon coding agents and million-token context work
Flagship GLM model for long-horizon coding, agents, and complex project delivery
Native multimodal GLM model for efficient coding and long-horizon agent tasks
Fast GLM vision model for screenshots, documents, and multimodal agent tasks
ZaiCodingPlan 系列
2 个模型xAI 系列
7 个模型Grok model for agentic tool use, reasoning, coding, and live assistance
Reasoning Grok for document-heavy analysis and long-horizon tool use
Grok model for agentic tool use, reasoning, coding, and live assistance
xAI's Grok for chat, coding, agentic tools, and lower hallucination risk
xAI's Grok model for chat, coding, agentic tools, and lower hallucination risk
xAI's frontier model for long-running agents, coding, knowledge work, and visual projects
Fast Grok coding model tuned for agentic engineering and iterative edits