Anthropic Messages API with 资产生成
资产生成 provides Anthropic Messages API endpoints, so you can use the Anthropic SDK and tools like Claude Code through a unified gateway with only a URL change.
The Anthropic Messages API implements the same specification as the Anthropic Messages API.
For more on using 资产生成 with Claude Code, see the Claude Code instructions.
The Anthropic Messages API is available at the following base URL:
https://ai-gateway.vercel.sh
The Anthropic Messages API supports the same authentication methods as the main 资产生成:
- API key: Use your 资产生成 API key with the
x-api-keyheader orAuthorization: Bearer <token>header - OIDC token: Use your Vercel OIDC token with the
Authorization: Bearer <token>header
You only need to use one of these forms of authentication. If an API key is specified it will take precedence over any OIDC token, even if the API key is invalid.
The 资产生成 supports the following Anthropic Messages API endpoints:
POST /v1/messages- Create messages, with support for streaming, tool calling, extended thinking, structured outputs, and imagesPOST /v1/messages/count_tokens- Count tokens in a message before sending it to Claude, for managing context windows and costs
POST /v1/messages/batches- Create a message batchGET /v1/messages/batches- List message batchesGET /v1/messages/batches/{batch_id}- Retrieve a message batchGET /v1/messages/batches/{batch_id}/results- Read message batch results
For advanced features, see:
- Extended thinking - Configure how much Claude thinks before answering
- Advanced features - Web search, provider timeouts, and automatic caching
- Structured outputs - JSON Schema-constrained responses
Claude Code is Anthropic's agentic coding tool. You can configure it to use AI 资产生成, enabling you to:
- Route requests through multiple AI providers
- Monitor traffic and spend in your 资产生成 Overview
- View detailed traces in Vercel Observability under AI
- Use any model available through the gateway
Configure Claude Code to use the 资产生成 by setting these environment variables:
Variable Value ANTHROPIC_BASE_URLhttps://ai-gateway.vercel.shANTHROPIC_AUTH_TOKENYour 资产生成 API key ANTHROPIC_API_KEY""(empty string)Add this alias to your
~/.zshrc(or~/.bashrc):alias claude-vercel='ANTHROPIC_BASE_URL="https://ai-gateway.vercel.sh" ANTHROPIC_AUTH_TOKEN="your-api-key-here" ANTHROPIC_API_KEY="" claude'Then reload your shell:
source ~/.zshrcFor more flexibility (e.g., adding additional logic), create a wrapper script at
~/bin/claude-vercel:claude-vercel#!/usr/bin/env bash # Routes Claude Code through AI 资产生成 ANTHROPIC_BASE_URL="https://ai-gateway.vercel.sh" \ ANTHROPIC_AUTH_TOKEN="your-api-key-here" \ ANTHROPIC_API_KEY="" \ claude "$@"Make it executable and ensure
~/binis in your PATH:mkdir -p ~/bin chmod +x ~/bin/claude-vercel echo 'export PATH="$HOME/bin:$PATH"' >> ~/.zshrc source ~/.zshrcRun
claude-vercelto start Claude Code with 资产生成:claude-vercelYour requests will now be routed through AI 资产生成.
You can use the 资产生成's Anthropic Messages API with the official Anthropic SDK. Point your client to the 资产生成's base URL and use your 资产生成 API key or OIDC token for authentication.
import Anthropic from '@anthropic-ai/sdk';
const anthropic = new Anthropic({
apiKey: process.env.AI_GATEWAY_API_KEY,
baseURL: 'https://ai-gateway.vercel.sh',
});
const message = await anthropic.messages.create({
model: 'anthropic/claude-opus-5',
max_tokens: 1024,
messages: [{ role: 'user', content: 'Hello, world!' }],
});import os
import anthropic
client = anthropic.Anthropic(
api_key=os.getenv('AI_GATEWAY_API_KEY'),
base_url='https://ai-gateway.vercel.sh'
)
message = client.messages.create(
model='anthropic/claude-opus-5',
max_tokens=1024,
messages=[
{'role': 'user', 'content': 'Hello, world!'}
]
)curl -X POST "https://ai-gateway.vercel.sh/v1/messages" \
-H "Authorization: Bearer $AI_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-opus-5",
"max_tokens": 1024,
"messages": [
{
"role": "user",
"content": "Hello, world!"
}
]
}'Use the Anthropic SDK's client.messages.batches methods to submit requests for asynchronous processing, poll for completion, and read results.
Install @anthropic-ai/sdk for TypeScript or JavaScript, or anthropic for Python. Set AI_GATEWAY_API_KEY to your 资产生成 API key. Use https://ai-gateway.vercel.sh as the SDK base URL, without /v1.
Set AI_GATEWAY_BATCH_IDEMPOTENCY_KEY to a value you persist for this batch, such as sentiment-eval-2026-09-17. Reuse that value and the identical payload when retrying submission. 资产生成 returns the existing batch for a matching retry. A new key creates a new batch; reusing a key with a different payload returns an error.
The example prints the batch ID, checks status up to 20 times at 30-second intervals, and reads results when processing ends:
import Anthropic from '@anthropic-ai/sdk';
import { setTimeout } from 'node:timers/promises';
const client = new Anthropic({
apiKey: process.env.AI_GATEWAY_API_KEY,
baseURL: 'https://ai-gateway.vercel.sh',
});
const key = process.env.AI_GATEWAY_BATCH_IDEMPOTENCY_KEY;
if (!key) throw new Error('Set AI_GATEWAY_BATCH_IDEMPOTENCY_KEY');
let batch = await client.messages.batches.create(
{
requests: [{
custom_id: 'review-1',
params: {
model: 'anthropic/claude-haiku-4.5',
max_tokens: 64,
messages: [{ role: 'user', content: 'Classify as positive or negative: I love this app.' }],
},
}],
},
{ headers: { 'Idempotency-Key': key } },
);
console.log('Batch ID:', batch.id);
for (let attempt = 0; attempt < 20 && batch.processing_status !== 'ended'; attempt++) {
await setTimeout(30_000);
batch = await client.messages.batches.retrieve(batch.id);
}
if (batch.processing_status !== 'ended') {
throw new Error(`Still processing. Resume polling batch ${batch.id} later.`);
}
if (!batch.results_url) {
throw new Error(`No downloadable results for batch ${batch.id}.`);
}
const results = await client.messages.batches.results(batch.id);
for await (const item of results) {
if (item.result.type === 'succeeded') {
console.log(item.custom_id, item.result.message.content);
} else {
console.error(item.custom_id, item.result);
}
}import Anthropic from '@anthropic-ai/sdk';
import { setTimeout } from 'node:timers/promises';
const client = new Anthropic({
apiKey: process.env.AI_GATEWAY_API_KEY,
baseURL: 'https://ai-gateway.vercel.sh',
});
const key = process.env.AI_GATEWAY_BATCH_IDEMPOTENCY_KEY;
if (!key) throw new Error('Set AI_GATEWAY_BATCH_IDEMPOTENCY_KEY');
let batch = await client.messages.batches.create(
{
requests: [{
custom_id: 'review-1',
params: {
model: 'anthropic/claude-haiku-4.5',
max_tokens: 64,
messages: [{ role: 'user', content: 'Classify as positive or negative: I love this app.' }],
},
}],
},
{ headers: { 'Idempotency-Key': key } },
);
console.log('Batch ID:', batch.id);
for (let attempt = 0; attempt < 20 && batch.processing_status !== 'ended'; attempt++) {
await setTimeout(30_000);
batch = await client.messages.batches.retrieve(batch.id);
}
if (batch.processing_status !== 'ended') {
throw new Error(`Still processing. Resume polling batch ${batch.id} later.`);
}
if (!batch.results_url) {
throw new Error(`No downloadable results for batch ${batch.id}.`);
}
const results = await client.messages.batches.results(batch.id);
for await (const item of results) {
if (item.result.type === 'succeeded') {
console.log(item.custom_id, item.result.message.content);
} else {
console.error(item.custom_id, item.result);
}
}import os
import time
import anthropic
client = anthropic.Anthropic(
api_key=os.environ['AI_GATEWAY_API_KEY'],
base_url='https://ai-gateway.vercel.sh',
)
batch = client.messages.batches.create(
requests=[{
'custom_id': 'review-1',
'params': {
'model': 'anthropic/claude-haiku-4.5',
'max_tokens': 64,
'messages': [{'role': 'user', 'content': 'Classify as positive or negative: I love this app.'}],
},
}],
extra_headers={'Idempotency-Key': os.environ['AI_GATEWAY_BATCH_IDEMPOTENCY_KEY']},
)
print('Batch ID:', batch.id, flush=True)
for _ in range(20):
if batch.processing_status == 'ended':
break
time.sleep(30)
batch = client.messages.batches.retrieve(batch.id)
if batch.processing_status != 'ended':
raise RuntimeError(f'Still processing. Resume polling batch {batch.id} later.')
if not batch.results_url:
raise RuntimeError(f'No downloadable results for batch {batch.id}.')
for item in client.messages.batches.results(batch.id):
if item.result.type == 'succeeded':
print(item.custom_id, item.result.message.content)
else:
print(item.custom_id, item.result)Save the printed batch ID to resume polling with client.messages.batches.retrieve later. The batch keeps processing after the example stops waiting.
When processing ends, a non-null results_url indicates that results are available. The SDK uses this URL when you call results. Individual requests can still fail. Match each result to its request by custom_id, regardless of result order. Successful results include result.message; other results report errored, canceled, or expired.
Use await client.messages.batches.list({ limit: 20 }) in TypeScript or JavaScript, or client.messages.batches.list(limit=20) in Python. The default page size is 20, with a maximum of 100. Pass either after_id or before_id to paginate, never both.
- Submit up to 1,000 requests per batch, all using the same model.
- Give each request a unique
custom_idof 1–64 characters using letters, numbers, underscores, or hyphens ([A-Za-z0-9_-]). - Batch requests don't support streaming, provider-executed tools, or Model Context Protocol (MCP) servers.
- Batch cancellation and deletion aren't supported through 资产生成.
- You can use Bring Your Own Key (BYOK) credentials saved to your team. Request-scoped raw provider keys aren't supported.
- Zero Data Retention (ZDR) isn't available for batches.
The messages endpoint supports the following parameters:
model(string): The model to use (e.g.,anthropic/claude-opus-5)max_tokens(integer): Maximum number of tokens to generatemessages(array): Array of message objects withroleandcontentfields
stream(boolean): Whether to stream the response. Defaults tofalsetemperature(number): Controls randomness in the output. Range: 0-1top_p(number): Nucleus sampling parameter. Range: 0-1top_k(integer): Top-k sampling parameterstop_sequences(array): Stop sequences for the generationtools(array): Array of tool definitions for function callingtool_choice(object): Controls which tools are calledthinking(object): Extended thinking configurationsystem(string or array): System prompt
The gateway passes through the cache_control parameter to Anthropic's prompt caching feature. This is explicit caching: you specify cache breakpoints, and Anthropic handles storing and reusing cached content automatically.
import fs from 'node:fs';
import Anthropic from '@anthropic-ai/sdk';
const apiKey = process.env.AI_GATEWAY_API_KEY || process.env.VERCEL_OIDC_TOKEN;
const anthropic = new Anthropic({
apiKey,
baseURL: 'https://ai-gateway.vercel.sh',
});
// Caching only pays off above a provider minimum, currently 1024 tokens for
// most Claude models. A short string is silently not cached.
const longDocumentContent = fs.readFileSync('./contract.txt', 'utf8');
const message = await anthropic.messages.create({
model: 'anthropic/claude-opus-5',
max_tokens: 1024,
system: [
{
type: 'text',
text: 'You are a helpful assistant that analyzes documents.',
},
{
type: 'text',
text: longDocumentContent,
cache_control: { type: 'ephemeral' },
},
],
messages: [
{
role: 'user',
content: 'Summarize the key points from this document.',
},
],
});
console.log(message.usage);
// {
// input_tokens: 50,
// output_tokens: 200,
// cache_creation_input_tokens: 10000, // Tokens written to cache
// cache_read_input_tokens: 0 // Tokens read from cache
// }import os
import anthropic
api_key = os.getenv('AI_GATEWAY_API_KEY') or os.getenv('VERCEL_OIDC_TOKEN')
client = anthropic.Anthropic(
api_key=api_key,
base_url='https://ai-gateway.vercel.sh'
)
# Caching only pays off above a provider minimum, currently 1024 tokens for
# most Claude models. A short string is silently not cached.
with open('contract.txt') as f:
long_document_content = f.read()
message = client.messages.create(
model='anthropic/claude-opus-5',
max_tokens=1024,
system=[
{
'type': 'text',
'text': 'You are a helpful assistant that analyzes documents.',
},
{
'type': 'text',
'text': long_document_content, # Large content to cache
'cache_control': {'type': 'ephemeral'},
},
],
messages=[
{
'role': 'user',
'content': 'Summarize the key points from this document.'
}
],
)
print(message.usage)
# {
# 'input_tokens': 50,
# 'output_tokens': 200,
# 'cache_creation_input_tokens': 10000, # Tokens written to cache
# 'cache_read_input_tokens': 0 # Tokens read from cache
# }CONTRACT=$(sed 's/"/\\"/g' contract.txt | tr '\n' ' ')
curl -X POST "https://ai-gateway.vercel.sh/v1/messages" \
-H "Authorization: Bearer $AI_GATEWAY_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "anthropic/claude-opus-5",
"max_tokens": 1024,
"system": [
{ "type": "text", "text": "You are a helpful assistant that analyzes documents." },
{
"type": "text",
"text": "'"$CONTRACT"'",
"cache_control": { "type": "ephemeral" }
}
],
"messages": [
{ "role": "user", "content": "Summarize the key points from this document." }
]
}'Add cache_control: { type: 'ephemeral' } to mark content that should be cached. You can place cache breakpoints on system messages, user message content, tool definitions, tool results, and assistant message content. Anthropic also supports automatic caching, where a single top-level cache_control field automatically applies to the last cacheable block.
For the full list of cacheable locations and automatic caching details, see the Anthropic prompt caching docs.
- First request: Content up to the breakpoint is cached (
cache_creation_input_tokens) - Subsequent requests: Matching prefixes are read from cache (
cache_read_input_tokens) - TTL: Cached content expires after 5 minutes, refreshed on each cache hit
The Claude Agent SDK (@anthropic-ai/claude-agent-sdk) lets you build agents with the same tools and agentic loop that power Claude Code. Because the SDK spawns Claude Code as a subprocess, it inherits the same ANTHROPIC_* environment variables described above, so your agent code needs no gateway-specific configuration:
import { query } from '@anthropic-ai/claude-agent-sdk';
for await (const message of query({
prompt: 'Find and fix the bug in auth.ts',
options: { allowedTools: ['Read', 'Edit', 'Bash'] },
})) {
console.log(message);
}Refer to the Claude Agent SDK documentation for more details.
The Agent SDK respects any environment variable the Claude Code CLI reads, including these two for working with 资产生成:
| Variable | Purpose |
|---|---|
CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS | Strips Anthropic-specific anthropic-beta headers and beta tool-schema fields from requests. Set to 1 when routing through providers like Bedrock or Vertex AI that reject those fields. |
CLAUDE_CODE_EXTRA_BODY | Merges a JSON object into the top level of every request body. Use it to pass providerOptions like order, only, and sort. |
For example, to restrict requests to Amazon Bedrock only, set these alongside the ANTHROPIC_* variables in your environment:
CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1
CLAUDE_CODE_EXTRA_BODY='{"providerOptions":{"gateway":{"only":["bedrock"]}}}'The API returns standard HTTP status codes and error responses:
400 Bad Request: Invalid request parameters401 Unauthorized: Invalid or missing authentication403 Forbidden: Insufficient permissions404 Not Found: Model or endpoint not found429 Too Many Requests: Rate limit exceeded500 Internal Server Error: Server error
{
"type": "error",
"error": {
"type": "invalid_request_error",
"message": "Invalid request: missing required parameter 'max_tokens'"
}
}Was this helpful?