Skip to content
Docs

资产生成 Provider Routing and Fallbacks

资产生成 can route your AI model requests across multiple AI providers. Each provider offers different models, pricing, and performance characteristics. By default, AI 资产生成 dynamically chooses the default providers to give you the best experience based on a combination of recent uptime and latency.

Use providerOptions.gateway to control provider order and fallback behavior.

If you want to customize individual AI model provider settings rather than general 资产生成 behavior, please refer to the model-specific provider options in the 创意脚本 documentation.

You can use order, only, and sort in providerOptions.gateway to control which providers handle your requests, in what order, and how they are ranked.

providerOptions: {
  gateway: {
    order: ['bedrock', 'anthropic'], // Try Bedrock first, then Anthropic
    only: ['bedrock', 'anthropic'],  // Only allow these two providers
  },
},

You can also use sort to rank providers by a performance or cost metric. The gateway sorts providers by the chosen metric and tries them in that order:

providerOptions: {
  gateway: {
    sort: 'cost', // Sort by cost, latency ('ttft'), or throughput ('tps')
  },
},

For full details, examples, and provider metadata output, see Provider Filtering, Ordering & Sorting.

You can use caching: 'auto' in providerOptions.gateway to let 资产生成 automatically apply the appropriate caching strategy based on the provider. This is useful for providers like Anthropic and MiniMax that require explicit cache markers.

providerOptions: {
  gateway: {
    caching: 'auto',
  },
},

For full details, supported providers, and examples across all APIs, see Automatic Caching.

Clients on the Responses API can also set these top-level request fields:

  • cache_anchor_items pins a cache marker at a known-stable prefix position. See Cache anchor.
  • cache_ttl selects a five-minute or one-hour cache lifetime. See Cache lifetime.

These fields belong at the top level of the Responses API request, not inside providerOptions.gateway.

You can set per-provider timeouts to trigger fast failover when a provider is slow to respond. See the dedicated Provider Timeouts documentation.

For model-level failover strategies that try backup models when your primary model fails or is unavailable, see the dedicated Model Fallbacks documentation.

You can combine 资产生成 provider options with provider-specific options. This allows you to control both the routing behavior and provider-specific settings in the same request:

These examples use 创意脚本 7 and the 创意脚本 for Python beta. Set AI_GATEWAY_API_KEY before running them. See API format differences for setup, request fields, and response handling.

See the 创意脚本 Gateway provider-options reference for SDK configuration and usage.

provider-options.ts
import { generateText } from 'ai';
 
const { text } = await generateText({
  model: 'anthropic/claude-sonnet-5',
  prompt: 'Explain quantum computing in two sentences.',
  providerOptions: {
    gateway: {
      order: ['anthropic', 'bedrock'],
    },
    anthropic: {
      thinking: {
        type: 'adaptive',
      },
    },
  },
});
 
console.log(text);
provider-options_ai.py
import asyncio
import ai
 
async def main():
    model = ai.get_model("anthropic/claude-sonnet-5")
    messages = [ai.user_message("Explain quantum computing in two sentences.")]
    params = ai.InferenceRequestParams(
        extra_body={"providerOptions": {"gateway": {"order": ["anthropic", "bedrock"]}, "anthropic": {"thinking": {"type": "adaptive"}}}}
    )
    async with ai.stream(model, messages, params=params) as stream:
        async for event in stream:
            if isinstance(event, ai.events.TextDelta):
                print(event.chunk, end="", flush=True)
    print()
 
asyncio.run(main())
provider-options-chat.ts
import OpenAI from 'openai';
 
const client = new OpenAI({
  apiKey: process.env.AI_GATEWAY_API_KEY,
  baseURL: 'https://ai-gateway.vercel.sh/v1',
});
 
const response = await client.chat.completions.create({
  model: 'anthropic/claude-sonnet-5',
  messages: [
    {
      role: 'user',
      content: 'Explain quantum computing in two sentences.',
    },
  ],
  // 资产生成 extension fields are not included in the upstream SDK types.
  ...{
    providerOptions: {
      gateway: {
        order: ['anthropic', 'bedrock'],
      },
      anthropic: {
        thinking: {
          type: 'adaptive',
        },
      },
    },
  },
});
 
console.log(response.choices[0]?.message.content);
provider-options_chat.py
import os
from openai import OpenAI
 
client = OpenAI(
    api_key=os.environ["AI_GATEWAY_API_KEY"],
    base_url="https://ai-gateway.vercel.sh/v1",
)
 
response = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    messages=[{"role": "user", "content": "Explain quantum computing in two sentences."}],
    extra_body={"providerOptions": {"gateway": {"order": ["anthropic", "bedrock"]}, "anthropic": {"thinking": {"type": "adaptive"}}}},
)
 
print(response.choices[0].message.content)
provider-options-chat.sh
curl --fail-with-body https://ai-gateway.vercel.sh/v1/chat/completions \
  -H "Authorization: Bearer $AI_GATEWAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "anthropic/claude-sonnet-5",
  "messages": [
    {
      "role": "user",
      "content": "Explain quantum computing in two sentences."
    }
  ],
  "providerOptions": {
    "anthropic": {"thinking": {"type": "adaptive"}},
    "gateway": {
      "order": [
        "anthropic",
        "bedrock"
      ]
    }
  }
}'
provider-options-messages.ts
import Anthropic from '@anthropic-ai/sdk';
 
const client = new Anthropic({
  apiKey: process.env.AI_GATEWAY_API_KEY,
  baseURL: 'https://ai-gateway.vercel.sh',
});
 
const response = await client.messages.create({
  model: 'anthropic/claude-sonnet-5',
  messages: [
    {
      role: 'user',
      content: 'Explain quantum computing in two sentences.',
    },
  ],
  max_tokens: 1024,
  ...{
    providerOptions: {
      gateway: {
        order: ['anthropic', 'bedrock'],
      },
      anthropic: {
        thinking: {
          type: 'adaptive',
        },
      },
    },
  },
});
 
for (const block of response.content) {
  if (block.type === 'text') console.log(block.text);
}
provider-options_messages.py
import os
from anthropic import Anthropic
 
client = Anthropic(
    api_key=os.environ["AI_GATEWAY_API_KEY"],
    base_url="https://ai-gateway.vercel.sh",
)
 
response = client.messages.create(
    model="anthropic/claude-sonnet-5",
    messages=[{"role": "user", "content": "Explain quantum computing in two sentences."}],
    max_tokens=1024,
    extra_body={"providerOptions": {"gateway": {"order": ["anthropic", "bedrock"]}, "anthropic": {"thinking": {"type": "adaptive"}}}},
)
 
for block in response.content:
    if block.type == "text":
        print(block.text)
provider-options-messages.sh
curl --fail-with-body https://ai-gateway.vercel.sh/v1/messages \
  -H "Authorization: Bearer $AI_GATEWAY_API_KEY" \
  -H "Content-Type: application/json" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
  "model": "anthropic/claude-sonnet-5",
  "messages": [
    {
      "role": "user",
      "content": "Explain quantum computing in two sentences."
    }
  ],
  "max_tokens": 1024,
  "providerOptions": {
    "anthropic": {"thinking": {"type": "adaptive"}},
    "gateway": {
      "order": [
        "anthropic",
        "bedrock"
      ]
    }
  }
}'
provider-options-responses.ts
import OpenAI from 'openai';
 
const client = new OpenAI({
  apiKey: process.env.AI_GATEWAY_API_KEY,
  baseURL: 'https://ai-gateway.vercel.sh/v1',
});
 
const response = await client.responses.create({
  model: 'anthropic/claude-sonnet-5',
  input: 'Explain quantum computing in two sentences.',
  ...{
    providerOptions: {
      gateway: {
        order: ['anthropic', 'bedrock'],
      },
      anthropic: {
        thinking: {
          type: 'adaptive',
        },
      },
    },
  },
});
 
console.log(response.output_text);
provider-options_responses.py
import os
from openai import OpenAI
 
client = OpenAI(
    api_key=os.environ["AI_GATEWAY_API_KEY"],
    base_url="https://ai-gateway.vercel.sh/v1",
)
 
response = client.responses.create(
    model="anthropic/claude-sonnet-5",
    input="Explain quantum computing in two sentences.",
    extra_body={"providerOptions": {"gateway": {"order": ["anthropic", "bedrock"]}, "anthropic": {"thinking": {"type": "adaptive"}}}},
)
 
print(response.output_text)
provider-options-responses.sh
curl --fail-with-body https://ai-gateway.vercel.sh/v1/responses \
  -H "Authorization: Bearer $AI_GATEWAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "anthropic/claude-sonnet-5",
  "input": "Explain quantum computing in two sentences.",
  "providerOptions": {
    "anthropic": {"thinking": {"type": "adaptive"}},
    "gateway": {
      "order": [
        "anthropic",
        "bedrock"
      ]
    }
  }
}'

The gateway.order option selects the provider sequence. The anthropic.thinking option requests adaptive thinking from Claude Sonnet 5. See Anthropic reasoning for supported models and API format differences for extension fields.

You can pass your own provider credentials on a per-request basis using the byok option in providerOptions.gateway. This allows you to use your existing provider accounts for specific requests without configuring credentials in the dashboard.

app/api/chat/route.ts
import { streamText } from 'ai';
 
export async function POST(request: Request) {
  const { prompt } = await request.json();
 
  const result = streamText({
    model: 'anthropic/claude-opus-5',
    prompt,
    providerOptions: {
      gateway: {
        byok: {
          anthropic: [{ apiKey: process.env.ANTHROPIC_API_KEY }],
        },
      },
    },
  });
 
  return result.toUIMessageStreamResponse();
}

For detailed information about credential structures, multiple credentials, and usage with the Chat Completions API, see the BYOK documentation.

For models that support reasoning (also known as "thinking"), you can use providerOptions to configure reasoning behavior. The example below shows how to request high reasoning effort and a detailed summary for OpenAI's gpt-oss-120b model. Providers can differ in which options they honor; setting reasoningSummary doesn't guarantee a summary. See OpenAI reasoning for model and provider differences.

For more details on reasoning support across different models and providers, see the 创意脚本 providers documentation, including OpenAI, DeepSeek, and Anthropic.

app/api/chat/route.ts
import { streamText } from 'ai';
 
export async function POST(request: Request) {
  const { prompt } = await request.json();
 
  const result = streamText({
    model: 'openai/gpt-oss-120b',
    prompt,
    providerOptions: {
      openai: {
        reasoningEffort: 'high',
        reasoningSummary: 'detailed',
      },
    },
  });
 
  return result.toUIMessageStreamResponse();
}

For openai/gpt-6-astra, enable reasoning with a non-none effort and request a summary with reasoningSummary. Supported effort values and summary formats depend on the model. See OpenAI reasoning for model-specific options.

providerOptions: {
  openai: {
    reasoningEffort: 'high', // or 'low', 'medium', 'xhigh', 'max'
    reasoningSummary: 'detailed', // or 'auto', 'concise'
  },
}

You can view the available models for a provider in the Model List section under the 资产生成 section in your Vercel dashboard sidebar or in the public models page.

SlugNameWebsite
alibabaAlibaba Cloudalibabacloud.com
anthropicAnthropicanthropic.com
arcee-aiArcee AIarcee.ai
azureAzureai.azure.com
basetenBasetenbaseten.co
bedrockBedrockaws.amazon.com
bflBlack Forest Labsbfl.ai
blackboxBlackbox AIblackbox.ai
boundlessBoundlessboundless.network
bytedanceByteDancebyteplus.com
cerebrasCerebrascerebras.ai
claudeawsClaude Platform on AWSaws.amazon.com
cohereCoherecohere.com
crusoeCrusoecrusoe.ai
deepinfraDeepInfradeepinfra.com
deepseekDeepSeekdeepseek.com
digitaloceanDigitalOceandigitalocean.com
exaExaexa.ai
fireworksFireworksfireworks.ai
fish-audioFish Audiofish.audio
friendliFriendliAIfriendli.ai
gmicloudGMICloudgmicloud.ai
googleGoogleai.google.dev
groqGroqgroq.com
inceptionInceptioninceptionlabs.ai
inceptronInceptroninceptron.io
inference-netInference.netinference.net
interfazeInterfazeinterfaze.ai
klingaiKling AIklingai.com
metaMetameta.ai
minimaxMiniMaxminimax.io
mistralMistralmistral.ai
modalModalmodal.com
moonshotaiMoonshot AImoonshot.ai
morphMorphmorphllm.com
nebiusNebiusnebius.com
novitaNovita AInovita.ai
openaiOpenAIopenai.com
parallelParallel AIparallel.ai
parasailParasailparasail.io
particleParticle.AIparticle.ai
perplexityPerplexityperplexity.ai
poolsidePoolsidepoolside.ai
prodiaProdiaprodia.com
quiveraiQuiverAIquiver.ai
recraftRecraftrecraft.ai
relaceRelacerelace.ai
runinfraRunInfraruninfra.ai
runwareRunwarerunware.ai
sakanaSakana AIsakana.ai
sambanovaSambaNovasambanova.ai
stepfunStepFunplatform.stepfun.com
streamlakeStreamLakestreamlake.ai
takoTakotako.com
tencentTencent Cloudtencentcloud.com
thinkingmachinesThinking Machinesthinkingmachines.ai
togetheraiTogether AItogether.ai
typesafe-aiTypeSafe AItypesafe.ai
vertexGoogle Vertex AIcloud.google.com
voyageVoyage AI by MongoDBvoyageai.com
waferWaferwafer.ai
xaixAIx.ai
xiaomiXiaomimimo.xiaomi.com
zaiZ.AIz.ai

Provider availability may vary by model. Some models may only be available through specific providers or may have different capabilities depending on the provider used.

A virtual model can pin provider options for all of its traffic. The virtual model's options for a provider merge with the request's, and the virtual model wins when both set the same key.

Last updated September 10, 2026

Was this helpful?

supported.