The AI Gateway for developers
never a cent more.
Invoicing is available with no payment processing fees. See the 资产生成 pricing docs.
Optimize routing for availability, cost, or latency
If a provider degrades, the Gateway fails over to the same model on another provider. Identical output, no downtime.
Everyday requests route to a cost-efficient open model. Complex jobs escalate to a frontier model only when needed.
Latency-sensitive requests route to the fastest-responding model for the lowest time to first token. Heavier requests fall through to a larger model.
Latency figures are illustrative and vary by region, traffic, and prompt length.
Routing, billing, and observability in one place
为创作者与团队提供云端视频 Agent 工作台:脚本、分镜、资产与成片一体。
- One API key, hundreds of models. Unified billing and observability across your entire AI stack, with text, image, video, and audio models.
- Route on behavior, fallback anytime. Automatic fallbacks during provider outages so your app stays up even when a model goes down.
- OpenAI$12.47Anthropic$8.22SpaceXAI$4.31Platform fee$0.00Total$25.00No markup, just fair prices. Pay exactly what providers charge with no platform fees.
Works with your existing AI stack
- X_AI_API_KEY- ANTHROPIC_API_KEY- OPEN_AI_API_KEY+ AI_GATEWAY_API_KEY在工作台内统一选用生图、生视频与配音能力,无需自己管理多套密钥。
Recent ships
“Moving to the gateway is just so ergonomic. We get references to model names, and rely on Vercel to do the correct implementations and handle the edge cases.”
- 20xImprovement in AI
reliability after switching - 30sTo adopt a new model,
down from an hour of code - 25%Improvement in average
latency on model calls
Security and compliance
Route only to ZDR providers, no training or prompt logging, configurable per request or team-wide.
- Zero Data RetentionEnforced on every request across your team. Routes only to providers under a ZDR agreement.
- No training on your dataRoute only to providers that will not train on customer data, configurable per request.
- Provider allowlistRestrict your team to approved providers. Enforced on every request, no code changes.
Take control
Manage usage across teams with budgets, quotas, and full request visibility.
Every modality you need
Text, image, video, realtime, speech, transcription, embeddings, and reranking through one endpoint.
Text
Image
Video
Realtime
Speech
Transcription
Embeddings
Reranking
OpenAI
Vercel is a cloud platform for deploying and hosting websites, apps, and serverless functions with speed, scalability, and simplicity. It gives teams one workflow from development to global delivery, so products ship faster while keeping reliability and performance high across every request. With robust developer tooling and seamless integrations, Vercel enables engineering teams to collaborate efficiently and manage code from preview to production. Its edge network and automatic scaling help you serve users globally with minimal latency and maximum uptime, so you can focus on building great products while Vercel handles infrastructure, deployments, and optimizations to ensure a fast, consistent user experience at scale.
[0.234, -0.198, 0.567, 0.012, -0.823, 0.445, -0.671, 0.298, 0.789, -0.234, 0.617, 0.082, -0.301, 0.519, -0.448, 0.176, 0.902, -0.057, 0.388, -0.741, 0.263, 0.014, -0.598, 0.831, -0.122, 0.476, 0.249, -0.687, 0.354, 0.911, -0.205, 0.066, 0.498, -0.379, 0.732, -0.461, 0.187, 0.853, -0.092, 0.541, 0.318, -0.226, 0.679, -0.518, 0.034, 0.793, -0.347, 0.612, 0.158, -0.806, 0.273, 0.469, -0.135, 0.587, 0.821, -0.044, 0.396, -0.752, 0.218, 0.503, -0.661, 0.129, 0.874, -0.317, 0.452, 0.085, -0.539, 0.706, -0.198, 0.361, 0.927, -0.475, 0.244, -0.683, 0.518, 0.039, -0.792, 0.155, 0.642, -0.288, 0.471, 0.836, -0.107, 0.523, -0.366, 0.018, 0.748, -0.491, 0.265, 0.882, -0.146, 0.397, 0.609, -0.728, 0.184, 0.456, -0.029, 0.713, -0.554, 0.298, 0.067, -0.421, 0.836, -0.193, 0.512, 0.347, -0.768, 0.124, 0.658, -0.385, 0.901, 0.071, -0.469, 0.234, 0.587, -0.812, 0.156, 0.473, -0.628, 0.319, 0.052, -0.741, 0.486, 0.207, -0.563, 0.894, -0.135, 0.428, 0.671, -0.298, 0.519, 0.084, -0.756, 0.347, 0.918, -0.412, 0.176, -0.689, 0.245, 0.561, -0.328, 0.073, 0.842, -0.197, 0.453, -0.726, 0.288, 0.617, 0.039, -0.504, 0.871, -0.265, 0.392, 0.158, -0.673, 0.527, 0.084, -0.439, 0.916, -0.221, 0.358, 0.495, -0.782, 0.146, 0.629, -0.317, 0.058, 0.847, -0.473, 0.196, 0.534, -0.628, 0.279, 0.913, -0.045, 0.461, 0.184, -0.752, 0.398, 0.625, -0.171, 0.043, 0.789, -0.526, 0.314, 0.867, -0.082, 0.471, -0.638, 0.219, 0.582, -0.395, 0.146, 0.704, -0.273, 0.519, 0.038, -0.846, 0.187, 0.493, -0.561, 0.328, 0.075, -0.412, 0.901, -0.246]
- 1Cache invalidation guide
- 2How streaming works
- 3Routing edge cases
- 4Provider failover patterns
Text
OpenAI
Vercel is a cloud platform for deploying and hosting websites, apps, and serverless functions with speed, scalability, and simplicity. It gives teams one workflow from development to global delivery, so products ship faster while keeping reliability and performance high across every request. With robust developer tooling and seamless integrations, Vercel enables engineering teams to collaborate efficiently and manage code from preview to production. Its edge network and automatic scaling help you serve users globally with minimal latency and maximum uptime, so you can focus on building great products while Vercel handles infrastructure, deployments, and optimizations to ensure a fast, consistent user experience at scale.
[0.234, -0.198, 0.567, 0.012, -0.823, 0.445, -0.671, 0.298, 0.789, -0.234, 0.617, 0.082, -0.301, 0.519, -0.448, 0.176, 0.902, -0.057, 0.388, -0.741, 0.263, 0.014, -0.598, 0.831, -0.122, 0.476, 0.249, -0.687, 0.354, 0.911, -0.205, 0.066, 0.498, -0.379, 0.732, -0.461, 0.187, 0.853, -0.092, 0.541, 0.318, -0.226, 0.679, -0.518, 0.034, 0.793, -0.347, 0.612, 0.158, -0.806, 0.273, 0.469, -0.135, 0.587, 0.821, -0.044, 0.396, -0.752, 0.218, 0.503, -0.661, 0.129, 0.874, -0.317, 0.452, 0.085, -0.539, 0.706, -0.198, 0.361, 0.927, -0.475, 0.244, -0.683, 0.518, 0.039, -0.792, 0.155, 0.642, -0.288, 0.471, 0.836, -0.107, 0.523, -0.366, 0.018, 0.748, -0.491, 0.265, 0.882, -0.146, 0.397, 0.609, -0.728, 0.184, 0.456, -0.029, 0.713, -0.554, 0.298, 0.067, -0.421, 0.836, -0.193, 0.512, 0.347, -0.768, 0.124, 0.658, -0.385, 0.901, 0.071, -0.469, 0.234, 0.587, -0.812, 0.156, 0.473, -0.628, 0.319, 0.052, -0.741, 0.486, 0.207, -0.563, 0.894, -0.135, 0.428, 0.671, -0.298, 0.519, 0.084, -0.756, 0.347, 0.918, -0.412, 0.176, -0.689, 0.245, 0.561, -0.328, 0.073, 0.842, -0.197, 0.453, -0.726, 0.288, 0.617, 0.039, -0.504, 0.871, -0.265, 0.392, 0.158, -0.673, 0.527, 0.084, -0.439, 0.916, -0.221, 0.358, 0.495, -0.782, 0.146, 0.629, -0.317, 0.058, 0.847, -0.473, 0.196, 0.534, -0.628, 0.279, 0.913, -0.045, 0.461, 0.184, -0.752, 0.398, 0.625, -0.171, 0.043, 0.789, -0.526, 0.314, 0.867, -0.082, 0.471, -0.638, 0.219, 0.582, -0.395, 0.146, 0.704, -0.273, 0.519, 0.038, -0.846, 0.187, 0.493, -0.561, 0.328, 0.075, -0.412, 0.901, -0.246]
- 1Cache invalidation guide
- 2How streaming works
- 3Routing edge cases
- 4Provider failover patterns
Image
Vercel is a cloud platform for deploying and hosting websites, apps, and serverless functions with speed, scalability, and simplicity. It gives teams one workflow from development to global delivery, so products ship faster while keeping reliability and performance high across every request. With robust developer tooling and seamless integrations, Vercel enables engineering teams to collaborate efficiently and manage code from preview to production. Its edge network and automatic scaling help you serve users globally with minimal latency and maximum uptime, so you can focus on building great products while Vercel handles infrastructure, deployments, and optimizations to ensure a fast, consistent user experience at scale.
[0.234, -0.198, 0.567, 0.012, -0.823, 0.445, -0.671, 0.298, 0.789, -0.234, 0.617, 0.082, -0.301, 0.519, -0.448, 0.176, 0.902, -0.057, 0.388, -0.741, 0.263, 0.014, -0.598, 0.831, -0.122, 0.476, 0.249, -0.687, 0.354, 0.911, -0.205, 0.066, 0.498, -0.379, 0.732, -0.461, 0.187, 0.853, -0.092, 0.541, 0.318, -0.226, 0.679, -0.518, 0.034, 0.793, -0.347, 0.612, 0.158, -0.806, 0.273, 0.469, -0.135, 0.587, 0.821, -0.044, 0.396, -0.752, 0.218, 0.503, -0.661, 0.129, 0.874, -0.317, 0.452, 0.085, -0.539, 0.706, -0.198, 0.361, 0.927, -0.475, 0.244, -0.683, 0.518, 0.039, -0.792, 0.155, 0.642, -0.288, 0.471, 0.836, -0.107, 0.523, -0.366, 0.018, 0.748, -0.491, 0.265, 0.882, -0.146, 0.397, 0.609, -0.728, 0.184, 0.456, -0.029, 0.713, -0.554, 0.298, 0.067, -0.421, 0.836, -0.193, 0.512, 0.347, -0.768, 0.124, 0.658, -0.385, 0.901, 0.071, -0.469, 0.234, 0.587, -0.812, 0.156, 0.473, -0.628, 0.319, 0.052, -0.741, 0.486, 0.207, -0.563, 0.894, -0.135, 0.428, 0.671, -0.298, 0.519, 0.084, -0.756, 0.347, 0.918, -0.412, 0.176, -0.689, 0.245, 0.561, -0.328, 0.073, 0.842, -0.197, 0.453, -0.726, 0.288, 0.617, 0.039, -0.504, 0.871, -0.265, 0.392, 0.158, -0.673, 0.527, 0.084, -0.439, 0.916, -0.221, 0.358, 0.495, -0.782, 0.146, 0.629, -0.317, 0.058, 0.847, -0.473, 0.196, 0.534, -0.628, 0.279, 0.913, -0.045, 0.461, 0.184, -0.752, 0.398, 0.625, -0.171, 0.043, 0.789, -0.526, 0.314, 0.867, -0.082, 0.471, -0.638, 0.219, 0.582, -0.395, 0.146, 0.704, -0.273, 0.519, 0.038, -0.846, 0.187, 0.493, -0.561, 0.328, 0.075, -0.412, 0.901, -0.246]
- 1Cache invalidation guide
- 2How streaming works
- 3Routing edge cases
- 4Provider failover patterns
Video
SpaceXAI
Vercel is a cloud platform for deploying and hosting websites, apps, and serverless functions with speed, scalability, and simplicity. It gives teams one workflow from development to global delivery, so products ship faster while keeping reliability and performance high across every request. With robust developer tooling and seamless integrations, Vercel enables engineering teams to collaborate efficiently and manage code from preview to production. Its edge network and automatic scaling help you serve users globally with minimal latency and maximum uptime, so you can focus on building great products while Vercel handles infrastructure, deployments, and optimizations to ensure a fast, consistent user experience at scale.
[0.234, -0.198, 0.567, 0.012, -0.823, 0.445, -0.671, 0.298, 0.789, -0.234, 0.617, 0.082, -0.301, 0.519, -0.448, 0.176, 0.902, -0.057, 0.388, -0.741, 0.263, 0.014, -0.598, 0.831, -0.122, 0.476, 0.249, -0.687, 0.354, 0.911, -0.205, 0.066, 0.498, -0.379, 0.732, -0.461, 0.187, 0.853, -0.092, 0.541, 0.318, -0.226, 0.679, -0.518, 0.034, 0.793, -0.347, 0.612, 0.158, -0.806, 0.273, 0.469, -0.135, 0.587, 0.821, -0.044, 0.396, -0.752, 0.218, 0.503, -0.661, 0.129, 0.874, -0.317, 0.452, 0.085, -0.539, 0.706, -0.198, 0.361, 0.927, -0.475, 0.244, -0.683, 0.518, 0.039, -0.792, 0.155, 0.642, -0.288, 0.471, 0.836, -0.107, 0.523, -0.366, 0.018, 0.748, -0.491, 0.265, 0.882, -0.146, 0.397, 0.609, -0.728, 0.184, 0.456, -0.029, 0.713, -0.554, 0.298, 0.067, -0.421, 0.836, -0.193, 0.512, 0.347, -0.768, 0.124, 0.658, -0.385, 0.901, 0.071, -0.469, 0.234, 0.587, -0.812, 0.156, 0.473, -0.628, 0.319, 0.052, -0.741, 0.486, 0.207, -0.563, 0.894, -0.135, 0.428, 0.671, -0.298, 0.519, 0.084, -0.756, 0.347, 0.918, -0.412, 0.176, -0.689, 0.245, 0.561, -0.328, 0.073, 0.842, -0.197, 0.453, -0.726, 0.288, 0.617, 0.039, -0.504, 0.871, -0.265, 0.392, 0.158, -0.673, 0.527, 0.084, -0.439, 0.916, -0.221, 0.358, 0.495, -0.782, 0.146, 0.629, -0.317, 0.058, 0.847, -0.473, 0.196, 0.534, -0.628, 0.279, 0.913, -0.045, 0.461, 0.184, -0.752, 0.398, 0.625, -0.171, 0.043, 0.789, -0.526, 0.314, 0.867, -0.082, 0.471, -0.638, 0.219, 0.582, -0.395, 0.146, 0.704, -0.273, 0.519, 0.038, -0.846, 0.187, 0.493, -0.561, 0.328, 0.075, -0.412, 0.901, -0.246]
- 1Cache invalidation guide
- 2How streaming works
- 3Routing edge cases
- 4Provider failover patterns
Realtime
OpenAI
Vercel is a cloud platform for deploying and hosting websites, apps, and serverless functions with speed, scalability, and simplicity. It gives teams one workflow from development to global delivery, so products ship faster while keeping reliability and performance high across every request. With robust developer tooling and seamless integrations, Vercel enables engineering teams to collaborate efficiently and manage code from preview to production. Its edge network and automatic scaling help you serve users globally with minimal latency and maximum uptime, so you can focus on building great products while Vercel handles infrastructure, deployments, and optimizations to ensure a fast, consistent user experience at scale.
[0.234, -0.198, 0.567, 0.012, -0.823, 0.445, -0.671, 0.298, 0.789, -0.234, 0.617, 0.082, -0.301, 0.519, -0.448, 0.176, 0.902, -0.057, 0.388, -0.741, 0.263, 0.014, -0.598, 0.831, -0.122, 0.476, 0.249, -0.687, 0.354, 0.911, -0.205, 0.066, 0.498, -0.379, 0.732, -0.461, 0.187, 0.853, -0.092, 0.541, 0.318, -0.226, 0.679, -0.518, 0.034, 0.793, -0.347, 0.612, 0.158, -0.806, 0.273, 0.469, -0.135, 0.587, 0.821, -0.044, 0.396, -0.752, 0.218, 0.503, -0.661, 0.129, 0.874, -0.317, 0.452, 0.085, -0.539, 0.706, -0.198, 0.361, 0.927, -0.475, 0.244, -0.683, 0.518, 0.039, -0.792, 0.155, 0.642, -0.288, 0.471, 0.836, -0.107, 0.523, -0.366, 0.018, 0.748, -0.491, 0.265, 0.882, -0.146, 0.397, 0.609, -0.728, 0.184, 0.456, -0.029, 0.713, -0.554, 0.298, 0.067, -0.421, 0.836, -0.193, 0.512, 0.347, -0.768, 0.124, 0.658, -0.385, 0.901, 0.071, -0.469, 0.234, 0.587, -0.812, 0.156, 0.473, -0.628, 0.319, 0.052, -0.741, 0.486, 0.207, -0.563, 0.894, -0.135, 0.428, 0.671, -0.298, 0.519, 0.084, -0.756, 0.347, 0.918, -0.412, 0.176, -0.689, 0.245, 0.561, -0.328, 0.073, 0.842, -0.197, 0.453, -0.726, 0.288, 0.617, 0.039, -0.504, 0.871, -0.265, 0.392, 0.158, -0.673, 0.527, 0.084, -0.439, 0.916, -0.221, 0.358, 0.495, -0.782, 0.146, 0.629, -0.317, 0.058, 0.847, -0.473, 0.196, 0.534, -0.628, 0.279, 0.913, -0.045, 0.461, 0.184, -0.752, 0.398, 0.625, -0.171, 0.043, 0.789, -0.526, 0.314, 0.867, -0.082, 0.471, -0.638, 0.219, 0.582, -0.395, 0.146, 0.704, -0.273, 0.519, 0.038, -0.846, 0.187, 0.493, -0.561, 0.328, 0.075, -0.412, 0.901, -0.246]
- 1Cache invalidation guide
- 2How streaming works
- 3Routing edge cases
- 4Provider failover patterns
Speech
ElevenLabs
Vercel is a cloud platform for deploying and hosting websites, apps, and serverless functions with speed, scalability, and simplicity. It gives teams one workflow from development to global delivery, so products ship faster while keeping reliability and performance high across every request. With robust developer tooling and seamless integrations, Vercel enables engineering teams to collaborate efficiently and manage code from preview to production. Its edge network and automatic scaling help you serve users globally with minimal latency and maximum uptime, so you can focus on building great products while Vercel handles infrastructure, deployments, and optimizations to ensure a fast, consistent user experience at scale.
[0.234, -0.198, 0.567, 0.012, -0.823, 0.445, -0.671, 0.298, 0.789, -0.234, 0.617, 0.082, -0.301, 0.519, -0.448, 0.176, 0.902, -0.057, 0.388, -0.741, 0.263, 0.014, -0.598, 0.831, -0.122, 0.476, 0.249, -0.687, 0.354, 0.911, -0.205, 0.066, 0.498, -0.379, 0.732, -0.461, 0.187, 0.853, -0.092, 0.541, 0.318, -0.226, 0.679, -0.518, 0.034, 0.793, -0.347, 0.612, 0.158, -0.806, 0.273, 0.469, -0.135, 0.587, 0.821, -0.044, 0.396, -0.752, 0.218, 0.503, -0.661, 0.129, 0.874, -0.317, 0.452, 0.085, -0.539, 0.706, -0.198, 0.361, 0.927, -0.475, 0.244, -0.683, 0.518, 0.039, -0.792, 0.155, 0.642, -0.288, 0.471, 0.836, -0.107, 0.523, -0.366, 0.018, 0.748, -0.491, 0.265, 0.882, -0.146, 0.397, 0.609, -0.728, 0.184, 0.456, -0.029, 0.713, -0.554, 0.298, 0.067, -0.421, 0.836, -0.193, 0.512, 0.347, -0.768, 0.124, 0.658, -0.385, 0.901, 0.071, -0.469, 0.234, 0.587, -0.812, 0.156, 0.473, -0.628, 0.319, 0.052, -0.741, 0.486, 0.207, -0.563, 0.894, -0.135, 0.428, 0.671, -0.298, 0.519, 0.084, -0.756, 0.347, 0.918, -0.412, 0.176, -0.689, 0.245, 0.561, -0.328, 0.073, 0.842, -0.197, 0.453, -0.726, 0.288, 0.617, 0.039, -0.504, 0.871, -0.265, 0.392, 0.158, -0.673, 0.527, 0.084, -0.439, 0.916, -0.221, 0.358, 0.495, -0.782, 0.146, 0.629, -0.317, 0.058, 0.847, -0.473, 0.196, 0.534, -0.628, 0.279, 0.913, -0.045, 0.461, 0.184, -0.752, 0.398, 0.625, -0.171, 0.043, 0.789, -0.526, 0.314, 0.867, -0.082, 0.471, -0.638, 0.219, 0.582, -0.395, 0.146, 0.704, -0.273, 0.519, 0.038, -0.846, 0.187, 0.493, -0.561, 0.328, 0.075, -0.412, 0.901, -0.246]
- 1Cache invalidation guide
- 2How streaming works
- 3Routing edge cases
- 4Provider failover patterns
Transcription
Deepgram
Vercel is a cloud platform for deploying and hosting websites, apps, and serverless functions with speed, scalability, and simplicity. It gives teams one workflow from development to global delivery, so products ship faster while keeping reliability and performance high across every request. With robust developer tooling and seamless integrations, Vercel enables engineering teams to collaborate efficiently and manage code from preview to production. Its edge network and automatic scaling help you serve users globally with minimal latency and maximum uptime, so you can focus on building great products while Vercel handles infrastructure, deployments, and optimizations to ensure a fast, consistent user experience at scale.
You can build and host many different types of applications from static sites with your favorite framework, multi-tenant applications or micro-frontends to AI-powered agents. Deploy globally in seconds, scale automatically with traffic, and ship every change with preview deployments, observability, and built-in security on every request.
[0.234, -0.198, 0.567, 0.012, -0.823, 0.445, -0.671, 0.298, 0.789, -0.234, 0.617, 0.082, -0.301, 0.519, -0.448, 0.176, 0.902, -0.057, 0.388, -0.741, 0.263, 0.014, -0.598, 0.831, -0.122, 0.476, 0.249, -0.687, 0.354, 0.911, -0.205, 0.066, 0.498, -0.379, 0.732, -0.461, 0.187, 0.853, -0.092, 0.541, 0.318, -0.226, 0.679, -0.518, 0.034, 0.793, -0.347, 0.612, 0.158, -0.806, 0.273, 0.469, -0.135, 0.587, 0.821, -0.044, 0.396, -0.752, 0.218, 0.503, -0.661, 0.129, 0.874, -0.317, 0.452, 0.085, -0.539, 0.706, -0.198, 0.361, 0.927, -0.475, 0.244, -0.683, 0.518, 0.039, -0.792, 0.155, 0.642, -0.288, 0.471, 0.836, -0.107, 0.523, -0.366, 0.018, 0.748, -0.491, 0.265, 0.882, -0.146, 0.397, 0.609, -0.728, 0.184, 0.456, -0.029, 0.713, -0.554, 0.298, 0.067, -0.421, 0.836, -0.193, 0.512, 0.347, -0.768, 0.124, 0.658, -0.385, 0.901, 0.071, -0.469, 0.234, 0.587, -0.812, 0.156, 0.473, -0.628, 0.319, 0.052, -0.741, 0.486, 0.207, -0.563, 0.894, -0.135, 0.428, 0.671, -0.298, 0.519, 0.084, -0.756, 0.347, 0.918, -0.412, 0.176, -0.689, 0.245, 0.561, -0.328, 0.073, 0.842, -0.197, 0.453, -0.726, 0.288, 0.617, 0.039, -0.504, 0.871, -0.265, 0.392, 0.158, -0.673, 0.527, 0.084, -0.439, 0.916, -0.221, 0.358, 0.495, -0.782, 0.146, 0.629, -0.317, 0.058, 0.847, -0.473, 0.196, 0.534, -0.628, 0.279, 0.913, -0.045, 0.461, 0.184, -0.752, 0.398, 0.625, -0.171, 0.043, 0.789, -0.526, 0.314, 0.867, -0.082, 0.471, -0.638, 0.219, 0.582, -0.395, 0.146, 0.704, -0.273, 0.519, 0.038, -0.846, 0.187, 0.493, -0.561, 0.328, 0.075, -0.412, 0.901, -0.246]
- 1Cache invalidation guide
- 2How streaming works
- 3Routing edge cases
- 4Provider failover patterns
Embeddings
OpenAI
Vercel is a cloud platform for deploying and hosting websites, apps, and serverless functions with speed, scalability, and simplicity. It gives teams one workflow from development to global delivery, so products ship faster while keeping reliability and performance high across every request. With robust developer tooling and seamless integrations, Vercel enables engineering teams to collaborate efficiently and manage code from preview to production. Its edge network and automatic scaling help you serve users globally with minimal latency and maximum uptime, so you can focus on building great products while Vercel handles infrastructure, deployments, and optimizations to ensure a fast, consistent user experience at scale.
[0.234, -0.198, 0.567, 0.012, -0.823, 0.445, -0.671, 0.298, 0.789, -0.234, 0.617, 0.082, -0.301, 0.519, -0.448, 0.176, 0.902, -0.057, 0.388, -0.741, 0.263, 0.014, -0.598, 0.831, -0.122, 0.476, 0.249, -0.687, 0.354, 0.911, -0.205, 0.066, 0.498, -0.379, 0.732, -0.461, 0.187, 0.853, -0.092, 0.541, 0.318, -0.226, 0.679, -0.518, 0.034, 0.793, -0.347, 0.612, 0.158, -0.806, 0.273, 0.469, -0.135, 0.587, 0.821, -0.044, 0.396, -0.752, 0.218, 0.503, -0.661, 0.129, 0.874, -0.317, 0.452, 0.085, -0.539, 0.706, -0.198, 0.361, 0.927, -0.475, 0.244, -0.683, 0.518, 0.039, -0.792, 0.155, 0.642, -0.288, 0.471, 0.836, -0.107, 0.523, -0.366, 0.018, 0.748, -0.491, 0.265, 0.882, -0.146, 0.397, 0.609, -0.728, 0.184, 0.456, -0.029, 0.713, -0.554, 0.298, 0.067, -0.421, 0.836, -0.193, 0.512, 0.347, -0.768, 0.124, 0.658, -0.385, 0.901, 0.071, -0.469, 0.234, 0.587, -0.812, 0.156, 0.473, -0.628, 0.319, 0.052, -0.741, 0.486, 0.207, -0.563, 0.894, -0.135, 0.428, 0.671, -0.298, 0.519, 0.084, -0.756, 0.347, 0.918, -0.412, 0.176, -0.689, 0.245, 0.561, -0.328, 0.073, 0.842, -0.197, 0.453, -0.726, 0.288, 0.617, 0.039, -0.504, 0.871, -0.265, 0.392, 0.158, -0.673, 0.527, 0.084, -0.439, 0.916, -0.221, 0.358, 0.495, -0.782, 0.146, 0.629, -0.317, 0.058, 0.847, -0.473, 0.196, 0.534, -0.628, 0.279, 0.913, -0.045, 0.461, 0.184, -0.752, 0.398, 0.625, -0.171, 0.043, 0.789, -0.526, 0.314, 0.867, -0.082, 0.471, -0.638, 0.219, 0.582, -0.395, 0.146, 0.704, -0.273, 0.519, 0.038, -0.846, 0.187, 0.493, -0.561, 0.328, 0.075, -0.412, 0.901, -0.246]
- 1Cache invalidation guide
- 2How streaming works
- 3Routing edge cases
- 4Provider failover patterns
Reranking
Cohere
Vercel is a cloud platform for deploying and hosting websites, apps, and serverless functions with speed, scalability, and simplicity. It gives teams one workflow from development to global delivery, so products ship faster while keeping reliability and performance high across every request. With robust developer tooling and seamless integrations, Vercel enables engineering teams to collaborate efficiently and manage code from preview to production. Its edge network and automatic scaling help you serve users globally with minimal latency and maximum uptime, so you can focus on building great products while Vercel handles infrastructure, deployments, and optimizations to ensure a fast, consistent user experience at scale.
[0.234, -0.198, 0.567, 0.012, -0.823, 0.445, -0.671, 0.298, 0.789, -0.234, 0.617, 0.082, -0.301, 0.519, -0.448, 0.176, 0.902, -0.057, 0.388, -0.741, 0.263, 0.014, -0.598, 0.831, -0.122, 0.476, 0.249, -0.687, 0.354, 0.911, -0.205, 0.066, 0.498, -0.379, 0.732, -0.461, 0.187, 0.853, -0.092, 0.541, 0.318, -0.226, 0.679, -0.518, 0.034, 0.793, -0.347, 0.612, 0.158, -0.806, 0.273, 0.469, -0.135, 0.587, 0.821, -0.044, 0.396, -0.752, 0.218, 0.503, -0.661, 0.129, 0.874, -0.317, 0.452, 0.085, -0.539, 0.706, -0.198, 0.361, 0.927, -0.475, 0.244, -0.683, 0.518, 0.039, -0.792, 0.155, 0.642, -0.288, 0.471, 0.836, -0.107, 0.523, -0.366, 0.018, 0.748, -0.491, 0.265, 0.882, -0.146, 0.397, 0.609, -0.728, 0.184, 0.456, -0.029, 0.713, -0.554, 0.298, 0.067, -0.421, 0.836, -0.193, 0.512, 0.347, -0.768, 0.124, 0.658, -0.385, 0.901, 0.071, -0.469, 0.234, 0.587, -0.812, 0.156, 0.473, -0.628, 0.319, 0.052, -0.741, 0.486, 0.207, -0.563, 0.894, -0.135, 0.428, 0.671, -0.298, 0.519, 0.084, -0.756, 0.347, 0.918, -0.412, 0.176, -0.689, 0.245, 0.561, -0.328, 0.073, 0.842, -0.197, 0.453, -0.726, 0.288, 0.617, 0.039, -0.504, 0.871, -0.265, 0.392, 0.158, -0.673, 0.527, 0.084, -0.439, 0.916, -0.221, 0.358, 0.495, -0.782, 0.146, 0.629, -0.317, 0.058, 0.847, -0.473, 0.196, 0.534, -0.628, 0.279, 0.913, -0.045, 0.461, 0.184, -0.752, 0.398, 0.625, -0.171, 0.043, 0.789, -0.526, 0.314, 0.867, -0.082, 0.471, -0.638, 0.219, 0.582, -0.395, 0.146, 0.704, -0.273, 0.519, 0.038, -0.846, 0.187, 0.493, -0.561, 0.328, 0.075, -0.412, 0.901, -0.246]
- 1Cache invalidation guide
- 2How streaming works
- 3Routing edge cases
- 4Provider failover patterns
Works with the tools your team already uses
Route 11+ AI coding agents through 资产生成 with a base URL change. Get unified observability and spend tracking across every tool, no matter who built it.
Or skip the config files. The setup command detects the supported agents on your machine, provisions an API key, and writes their configuration for you.
配置资产生成- Claude CodeAnthropic’s coding agent. Route it through 资产生成’s Anthropic-compatible endpoint, then switch between Claude models mid-session. Works with Claude Code Max too.
Switch models with
/model - OpenAI CodexOpenAI’s coding agent. Route it through 资产生成’s Responses API so usage joins the rest of your AI spend, and pick an OpenAI model per session.
Switch models with
/model - Command CodeTerminal coding agent with bring your own key support. Add your 资产生成 key and a custom base URL, then switch between your configured models.
Switch models with
/model - OpenCodeOpen-source terminal coding agent with native 资产生成 support. Connect once, then switch between any model on the fly.
Switch models with
/models - Blackbox AITerminal CLI for AI code generation and debugging, with access to every model in the catalog.
- ClineAutonomous coding agent for VS Code. Select AI 资产生成 as the provider for detailed token and cache metrics.
Get started
This quickstart walks you through making your first text generation request with 资产生成.
Create a new directory and initialize a Node.js project.
mkdir ai-text-democd ai-text-demopnpm init打开工作台即可开始创作,无需本地安装。
npm install ai@latest dotenv @types/node tsx typescriptGo to the 资产生成 API Keys page 在控制台中打开并点击 Create Key to generate a new API Key. Create a .env.local file and save your API Key. Instead of using an API Key, you can use OIDC tokens to authenticate your requests.
AI_GATEWAY_API_KEY=your_ai_gateway_api_keyCreate the index.ts file.
import { streamText } from 'ai';import 'dotenv/config';
async function main() { const result = streamText({ model: 'openai/gpt-5.5', prompt: 'Invent a new holiday and describe its traditions.', });
for await (const textPart of result.textStream) { process.stdout.write(textPart); }
console.log(); console.log('Token usage:', await result.usage); console.log('Finish reason:', await result.finishReason);}
main().catch(console.error);You should see the AI model’s response stream to your terminal.
pnpm tsx index.tsCreate a new directory and initialize a Node.js project.
mkdir ai-text-democd ai-text-demopnpm init打开工作台即可开始创作,无需本地安装。
npm install ai@latest dotenv @types/node tsx typescriptGo to the 资产生成 API Keys page 在控制台中打开并点击 Create Key to generate a new API Key. Create a .env.local file and save your API Key. Instead of using an API Key, you can use OIDC tokens to authenticate your requests.
AI_GATEWAY_API_KEY=your_ai_gateway_api_keyCreate the index.ts file.
import { streamText } from 'ai';import 'dotenv/config';
async function main() { const result = streamText({ model: 'openai/gpt-5.5', prompt: 'Invent a new holiday and describe its traditions.', });
for await (const textPart of result.textStream) { process.stdout.write(textPart); }
console.log(); console.log('Token usage:', await result.usage); console.log('Finish reason:', await result.finishReason);}
main().catch(console.error);You should see the AI model’s response stream to your terminal.
pnpm tsx index.tsFrequently asked questions
How is 资产生成 priced?
We offer tokens at list price from the upstream providers with no markup, including when you bring your own keys. Certain capabilities are available at higher plan tiers and metered separately. Invoicing is available with no payment processing fees. See the pricing page for details.
What's the difference between using 资产生成 and going direct to each provider?
资产生成 gives you one integration, automatic failover, unified spend tracking, and one invoice across every major provider. Going direct means signing N contracts and stitching together N billing dashboards.
Will 资产生成 work with our existing AI stack?
Almost certainly. 资产生成 supports the 创意脚本, OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, and an OpenResponses-compatible endpoint. Migrating is typically a base URL swap with no code changes.
Which modalities does 资产生成 support?
Text, image, video, embeddings, and reranking, all through the same endpoint. Browse the full model catalog for specific models and providers.
How does 资产生成 handle our enterprise security and compliance requirements?
资产生成 supports Zero Data Retention routing, a no-training guarantee, and team-wide provider allowlists. See the security overview for full details.
What observability does 资产生成 provide out of the box?
A dashboard with usage, spend, request volume, TTFT, and token counts, broken down by model, provider, and project. For deeper analysis, the Custom Reporting API lets you pull the same data into your own tools.
Can we use our existing provider contracts and committed spend?
Yes, through BYOK. Bring your own keys for almost every supported provider and your existing commitments flow through. We try BYOK first and only fall back to system credentials on failure.
Do I pay per request or get invoiced?
资产生成 uses pre-purchased credits by default. Top up in the dashboard and usage is drawn down per request. Enterprise customers can switch to a single consolidated invoice from Vercel covering every provider in their routing pool. For invoicing, reach out to sales for more details.
Can we purchase 资产生成 through AWS Marketplace?
Yes. 资产生成 is live on AWS Marketplace and available to purchase via AWS private offers. Procure it directly through AWS Marketplace and apply the spend toward your existing AWS cloud commits, within your existing budgets and procurement processes.