Skip to content
Docs

资产生成 Custom Reporting API

The Custom Reporting API gives you detailed visibility into your 资产生成 usage. You can break down costs and token consumption by model, user, tag, provider, or credential type to understand exactly where your AI spend is going.

Use it to:

  • Track costs by model: See how much you're spending on each model and compare cost efficiency across providers
  • Monitor per-user usage: Identify which users are driving the most spend and token consumption
  • Analyze by tags: Tag requests by feature, environment, or team to attribute costs and track usage across your organization
  • Compare providers: Understand cost and usage differences between providers serving the same models
  • Audit BYOK vs system credentials: Break down usage by credential type to see the impact of bring-your-own-key requests

The API is currently scoped to your entire account, so the API key you use will return usage data for everything on the account.

Charge typeCost
Write$0.075 / 1,000 tag/user ID writes
Query$5 / 1,000 queries to the reporting endpoint

Each unique tag or user ID within a single request scope counts as one write.

To use reporting, attach a user and/or tags to your 资产生成 requests. You can do this through the 创意脚本, Chat Completions API, Responses API, OpenResponses API, or Anthropic Messages API. For Chat Completions, the standard user field supplies the reporting user when providerOptions.gateway.user is not set.

These examples use 创意脚本 7 and the 创意脚本 for Python beta. Set AI_GATEWAY_API_KEY before running them. See API format differences for setup, request fields, and response handling.

See the 创意脚本 usage-tracking reference for SDK configuration and usage.

custom-reporting.ts
import { generateText } from 'ai';
 
const { text } = await generateText({
  model: 'anthropic/claude-sonnet-5',
  prompt: 'Tell me about San Francisco.',
  providerOptions: {
    gateway: {
      user: 'user-123',
      tags: ['feature:chat', 'env:development'],
    },
  },
});
 
console.log(text);
custom-reporting_ai.py
import asyncio
import ai
 
async def main():
    model = ai.get_model("anthropic/claude-sonnet-5")
    messages = [ai.user_message("Tell me about San Francisco.")]
    params = ai.InferenceRequestParams(
        extra_body={"providerOptions": {"gateway": {"user": "user-123", "tags": ["feature:chat", "env:development"]}}}
    )
    async with ai.stream(model, messages, params=params) as stream:
        async for event in stream:
            if isinstance(event, ai.events.TextDelta):
                print(event.chunk, end="", flush=True)
    print()
 
asyncio.run(main())
custom-reporting-chat.ts
import OpenAI from 'openai';
 
const client = new OpenAI({
  apiKey: process.env.AI_GATEWAY_API_KEY,
  baseURL: 'https://ai-gateway.vercel.sh/v1',
});
 
const response = await client.chat.completions.create({
  model: 'anthropic/claude-sonnet-5',
  messages: [
    {
      role: 'user',
      content: 'Tell me about San Francisco.',
    },
  ],
  // 资产生成 extension fields are not included in the upstream SDK types.
  ...{
    providerOptions: {
      gateway: {
        user: 'user-123',
        tags: ['feature:chat', 'env:development'],
      },
    },
  },
});
 
console.log(response.choices[0]?.message.content);
custom-reporting_chat.py
import os
from openai import OpenAI
 
client = OpenAI(
    api_key=os.environ["AI_GATEWAY_API_KEY"],
    base_url="https://ai-gateway.vercel.sh/v1",
)
 
response = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    messages=[{"role": "user", "content": "Tell me about San Francisco."}],
    extra_body={"providerOptions": {"gateway": {"user": "user-123", "tags": ["feature:chat", "env:development"]}}},
)
 
print(response.choices[0].message.content)
custom-reporting-chat.sh
curl --fail-with-body https://ai-gateway.vercel.sh/v1/chat/completions \
  -H "Authorization: Bearer $AI_GATEWAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "anthropic/claude-sonnet-5",
  "messages": [
    {
      "role": "user",
      "content": "Tell me about San Francisco."
    }
  ],
  "providerOptions": {
    "gateway": {
      "user": "user-123",
      "tags": [
        "feature:chat",
        "env:development"
      ]
    }
  }
}'
custom-reporting-messages.ts
import Anthropic from '@anthropic-ai/sdk';
 
const client = new Anthropic({
  apiKey: process.env.AI_GATEWAY_API_KEY,
  baseURL: 'https://ai-gateway.vercel.sh',
});
 
const response = await client.messages.create({
  model: 'anthropic/claude-sonnet-5',
  messages: [
    {
      role: 'user',
      content: 'Tell me about San Francisco.',
    },
  ],
  max_tokens: 1024,
  ...{
    providerOptions: {
      gateway: {
        user: 'user-123',
        tags: ['feature:chat', 'env:development'],
      },
    },
  },
});
 
for (const block of response.content) {
  if (block.type === 'text') console.log(block.text);
}
custom-reporting_messages.py
import os
from anthropic import Anthropic
 
client = Anthropic(
    api_key=os.environ["AI_GATEWAY_API_KEY"],
    base_url="https://ai-gateway.vercel.sh",
)
 
response = client.messages.create(
    model="anthropic/claude-sonnet-5",
    messages=[{"role": "user", "content": "Tell me about San Francisco."}],
    max_tokens=1024,
    extra_body={"providerOptions": {"gateway": {"user": "user-123", "tags": ["feature:chat", "env:development"]}}},
)
 
for block in response.content:
    if block.type == "text":
        print(block.text)
custom-reporting-messages.sh
curl --fail-with-body https://ai-gateway.vercel.sh/v1/messages \
  -H "Authorization: Bearer $AI_GATEWAY_API_KEY" \
  -H "Content-Type: application/json" \
  -H "anthropic-version: 2023-06-01" \
  -d '{
  "model": "anthropic/claude-sonnet-5",
  "messages": [
    {
      "role": "user",
      "content": "Tell me about San Francisco."
    }
  ],
  "max_tokens": 1024,
  "providerOptions": {
    "gateway": {
      "user": "user-123",
      "tags": [
        "feature:chat",
        "env:development"
      ]
    }
  }
}'
custom-reporting-responses.ts
import OpenAI from 'openai';
 
const client = new OpenAI({
  apiKey: process.env.AI_GATEWAY_API_KEY,
  baseURL: 'https://ai-gateway.vercel.sh/v1',
});
 
const response = await client.responses.create({
  model: 'anthropic/claude-sonnet-5',
  input: 'Tell me about San Francisco.',
  ...{
    providerOptions: {
      gateway: {
        user: 'user-123',
        tags: ['feature:chat', 'env:development'],
      },
    },
  },
});
 
console.log(response.output_text);
custom-reporting_responses.py
import os
from openai import OpenAI
 
client = OpenAI(
    api_key=os.environ["AI_GATEWAY_API_KEY"],
    base_url="https://ai-gateway.vercel.sh/v1",
)
 
response = client.responses.create(
    model="anthropic/claude-sonnet-5",
    input="Tell me about San Francisco.",
    extra_body={"providerOptions": {"gateway": {"user": "user-123", "tags": ["feature:chat", "env:development"]}}},
)
 
print(response.output_text)
custom-reporting-responses.sh
curl --fail-with-body https://ai-gateway.vercel.sh/v1/responses \
  -H "Authorization: Bearer $AI_GATEWAY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "anthropic/claude-sonnet-5",
  "input": "Tell me about San Francisco.",
  "providerOptions": {
    "gateway": {
      "user": "user-123",
      "tags": [
        "feature:chat",
        "env:development"
      ]
    }
  }
}'

You can also send reporting metadata as HTTP headers instead of (or in addition to) providerOptions.gateway. This is useful when a platform or proxy layer stamps context onto traffic without modifying application code:

HeaderTypeBehavior when the request body also sets the same field
ai-reporting-tagsstringComma-separated list. Merged with providerOptions.gateway.tags (deduped union).
ai-reporting-userstringSingle value. Overwrites providerOptions.gateway.user when present.

Validation limits match the body schema: up to 10 tags total after merging header and body values (deduped), with each tag between 1 and 64 characters; user up to 256 characters. An invalid header returns HTTP 400.

Both headers work across 资产生成 endpoints that accept providerOptions.gateway, including the formats shown below. The defaultHeaders / default_headers pattern on the SDK client is the same regardless of which endpoint you call. Swap in responses.create, messages.create, embeddings, image generation, or other supported calls as needed.

reporting-headers.ts
import { generateText } from 'ai';
 
const { text } = await generateText({
  model: 'anthropic/claude-sonnet-5',
  prompt: 'Explain quantum computing in two sentences.',
  headers: {
    'ai-reporting-tags': 'team:billing,feature:chat,env:development',
    'ai-reporting-user': 'user-12345',
  },
});
 
console.log(text);
reporting-headers_ai.py
import asyncio
import ai
 
async def main():
    model = ai.get_model("anthropic/claude-sonnet-5")
    messages = [ai.user_message("Explain quantum computing in two sentences.")]
    params = ai.InferenceRequestParams(
        extra_headers={"ai-reporting-tags": "team:billing,feature:chat,env:development", "ai-reporting-user": "user-12345"}
    )
    async with ai.stream(model, messages, params=params) as stream:
        async for event in stream:
            if isinstance(event, ai.events.TextDelta):
                print(event.chunk, end="", flush=True)
    print()
 
asyncio.run(main())
reporting-headers-chat.ts
import OpenAI from 'openai';
 
const client = new OpenAI({
  apiKey: process.env.AI_GATEWAY_API_KEY,
  baseURL: 'https://ai-gateway.vercel.sh/v1',
  defaultHeaders: {
    'ai-reporting-tags': 'team:billing,feature:chat,env:development',
    'ai-reporting-user': 'user-12345',
  },
});
 
const response = await client.chat.completions.create({
  model: 'anthropic/claude-sonnet-5',
  messages: [
    {
      role: 'user',
      content: 'Explain quantum computing in two sentences.',
    },
  ],
});
 
console.log(response.choices[0]?.message.content);
reporting-headers_chat.py
import os
from openai import OpenAI
 
client = OpenAI(
    api_key=os.environ["AI_GATEWAY_API_KEY"],
    base_url="https://ai-gateway.vercel.sh/v1",
    default_headers={"ai-reporting-tags": "team:billing,feature:chat,env:development", "ai-reporting-user": "user-12345"},
)
 
response = client.chat.completions.create(
    model="anthropic/claude-sonnet-5",
    messages=[{"role": "user", "content": "Explain quantum computing in two sentences."}],
)
 
print(response.choices[0].message.content)
reporting-headers-chat.sh
curl --fail-with-body https://ai-gateway.vercel.sh/v1/chat/completions \
  -H "Authorization: Bearer $AI_GATEWAY_API_KEY" \
  -H "Content-Type: application/json" \
  -H "ai-reporting-tags: team:billing,feature:chat,env:development" \
  -H "ai-reporting-user: user-12345" \
  -d '{
  "model": "anthropic/claude-sonnet-5",
  "messages": [
    {
      "role": "user",
      "content": "Explain quantum computing in two sentences."
    }
  ]
}'
reporting-headers-messages.ts
import Anthropic from '@anthropic-ai/sdk';
 
const client = new Anthropic({
  apiKey: process.env.AI_GATEWAY_API_KEY,
  baseURL: 'https://ai-gateway.vercel.sh',
  defaultHeaders: {
    'ai-reporting-tags': 'team:billing,feature:chat,env:development',
    'ai-reporting-user': 'user-12345',
  },
});
 
const response = await client.messages.create({
  model: 'anthropic/claude-sonnet-5',
  messages: [
    {
      role: 'user',
      content: 'Explain quantum computing in two sentences.',
    },
  ],
  max_tokens: 1024,
});
 
for (const block of response.content) {
  if (block.type === 'text') console.log(block.text);
}
reporting-headers_messages.py
import os
from anthropic import Anthropic
 
client = Anthropic(
    api_key=os.environ["AI_GATEWAY_API_KEY"],
    base_url="https://ai-gateway.vercel.sh",
    default_headers={"ai-reporting-tags": "team:billing,feature:chat,env:development", "ai-reporting-user": "user-12345"},
)
 
response = client.messages.create(
    model="anthropic/claude-sonnet-5",
    messages=[{"role": "user", "content": "Explain quantum computing in two sentences."}],
    max_tokens=1024,
)
 
for block in response.content:
    if block.type == "text":
        print(block.text)
reporting-headers-messages.sh
curl --fail-with-body https://ai-gateway.vercel.sh/v1/messages \
  -H "Authorization: Bearer $AI_GATEWAY_API_KEY" \
  -H "Content-Type: application/json" \
  -H "anthropic-version: 2023-06-01" \
  -H "ai-reporting-tags: team:billing,feature:chat,env:development" \
  -H "ai-reporting-user: user-12345" \
  -d '{
  "model": "anthropic/claude-sonnet-5",
  "messages": [
    {
      "role": "user",
      "content": "Explain quantum computing in two sentences."
    }
  ],
  "max_tokens": 1024
}'
reporting-headers-responses.ts
import OpenAI from 'openai';
 
const client = new OpenAI({
  apiKey: process.env.AI_GATEWAY_API_KEY,
  baseURL: 'https://ai-gateway.vercel.sh/v1',
  defaultHeaders: {
    'ai-reporting-tags': 'team:billing,feature:chat,env:development',
    'ai-reporting-user': 'user-12345',
  },
});
 
const response = await client.responses.create({
  model: 'anthropic/claude-sonnet-5',
  input: 'Explain quantum computing in two sentences.',
});
 
console.log(response.output_text);
reporting-headers_responses.py
import os
from openai import OpenAI
 
client = OpenAI(
    api_key=os.environ["AI_GATEWAY_API_KEY"],
    base_url="https://ai-gateway.vercel.sh/v1",
    default_headers={"ai-reporting-tags": "team:billing,feature:chat,env:development", "ai-reporting-user": "user-12345"},
)
 
response = client.responses.create(
    model="anthropic/claude-sonnet-5",
    input="Explain quantum computing in two sentences.",
)
 
print(response.output_text)
reporting-headers-responses.sh
curl --fail-with-body https://ai-gateway.vercel.sh/v1/responses \
  -H "Authorization: Bearer $AI_GATEWAY_API_KEY" \
  -H "Content-Type: application/json" \
  -H "ai-reporting-tags: team:billing,feature:chat,env:development" \
  -H "ai-reporting-user: user-12345" \
  -d '{
  "model": "anthropic/claude-sonnet-5",
  "input": "Explain quantum computing in two sentences."
}'

The reporting endpoint is available on Pro and Enterprise plans. The team is inferred from the API key or OIDC token. Hobby and Pro-trial plans cannot use this endpoint.

Endpoint
GET https://ai-gateway.vercel.sh/v1/report

All requests require a Bearer token in the Authorization header:

Authorization: Bearer YOUR_API_KEY
ParameterTypeDescription
start_datestringStart date in YYYY-MM-DD format
end_datestringEnd date in YYYY-MM-DD format

Dates are inclusive (both start_date and end_date are included) and in UTC.

ParameterTypeOptionsDescription
group_bystringday (default), user, model, tag, provider, credential_type, zero_data_retention, api_key_nameHow to aggregate the results. Each row represents one bucket of this dimension.
date_partstringday (default), hourTime granularity. Only applies when group_by=day. Use hour for per-hour rows, day for per-day rows.

Filters are applied before aggregation. Combine them with any group_by value.

ParameterTypeDescriptionExample
api_key_idstringFilter by a stable API key ID. Use self for the 资产生成 API key that authenticated the report request.abc123 or self
user_idstringFilter by a specific user IDuser_123
modelstringFilter by a specific model in creator/model-name formatanthropic/claude-sonnet-5
providerstringFilter by provideropenai
credential_typestringFilter by credential typebyok or system
zero_data_retentionbooleanFilter to Zero Data Retention (ZDR)-requested vs non-ZDR requeststrue or false
tagsstringFilter by one or more comma-separated tags. By default, requests match when they contain any listed tag.production or production,api
tags_matchstringMatch mode for tags. Use any to match requests with any listed tag, or all to require every listed tag. Defaults to any.any or all

API key names are not unique, so use the stable key ID when filtering. List the team's API keys to find each key's id. If you omit api_key_id, the report includes spend across the team. self requires 资产生成 API key authentication. The API returns a 400 response if you use self with an OIDC token, personal access token, or app token.

terminal
curl "https://ai-gateway.vercel.sh/v1/report?start_date=2026-01-01&end_date=2026-01-31&group_by=model" \
  -H "Authorization: Bearer $AI_GATEWAY_API_KEY"

The API returns a JSON object with a results array. Each row contains the one grouping field that matches the group_by parameter you used, plus the aggregated metrics. The example below shows every possible field together so you can see the shape; in a real response, only the grouping field for your selected group_by will be present. It can take a few minutes for requests to appear in the reporting endpoint.

Response
{
  "results": [
    {
      "day": "2026-01-01",
      "model": "anthropic/claude-sonnet-5",
      "provider": "anthropic",
      "user": "user_123",
      "tag": "production",
      "credential_type": "system",
      "zero_data_retention": "false",
      "api_key_name": "Production key",
      "total_cost": 10.5,
      "market_cost": 12.0,
      "surcharge_cost": 0.5,
      "gateway_cost": 0,
      "input_tokens": 1000,
      "output_tokens": 500,
      "cached_input_tokens": 200,
      "cache_creation_input_tokens": 50,
      "reasoning_tokens": 100,
      "request_count": 25
    }
  ]
}

Every row includes a single grouping field that depends on group_by, plus the metrics below.

FieldPresent whenTypeNotes
daygroup_by=day and date_part=day (default)stringThe UTC date for the bucket (YYYY-MM-DD)
hourgroup_by=day and date_part=hourstringThe UTC hour for the bucket (YYYY-MM-DDTHH)
usergroup_by=userstringThe user ID attached to the request
modelgroup_by=modelstringThe model in creator/model-name form
taggroup_by=tagstringA single tag value (one row per tag in the request)
providergroup_by=providerstringThe provider that served the request
credential_typegroup_by=credential_typestringbyok or system
zero_data_retentiongroup_by=zero_data_retentionstringtrue or false
api_key_namegroup_by=api_key_namestringThe human-readable name of the API key that served the request
FieldTypeDescription
total_costnumberCharged price in USD. Returns 0.00 for BYOK requests.
market_costnumberMarket price of the request at the time it ran. Includes both BYOK and non-BYOK cost.
surcharge_costnumberSurcharge portion of total_cost (for example, from add-on capabilities).
gateway_costnumber资产生成's own cost, separate from the provider rate.
input_tokensnumberInput tokens used
output_tokensnumberOutput tokens used
cached_input_tokensnumberCached input tokens
cache_creation_input_tokensnumberCache creation tokens
reasoning_tokensnumberReasoning tokens
request_countnumberNumber of requests in this row

All cost values are in USD and aggregated based on the grouping parameter.

Query spend reports with the 创意脚本's getSpendReport() method. It accepts the same parameters as the REST API (in camelCase) and returns camelCase results.

import { gateway } from 'ai';
 
const report = await gateway.getSpendReport({
  startDate: '2026-03-01',
  endDate: '2026-03-25',
  groupBy: 'model',
});
 
for (const row of report.results) {
  console.log(`${row.model}: $${row.totalCost.toFixed(4)}`);
}

You can combine tagging on requests with filtered queries to attribute costs by feature, team, or environment:

import type { GatewayProviderOptions } from '@ai-sdk/gateway';
import { gateway, streamText } from 'ai';
 
// 1. Make requests with tags
const result = streamText({
  model: 'anthropic/claude-opus-5',
  prompt: "Summarize this quarter's results",
  providerOptions: {
    gateway: {
      tags: ['team:finance', 'feature:summaries'],
    } satisfies GatewayProviderOptions,
  },
});
 
// 2. Later, query spend filtered by those tags
const report = await gateway.getSpendReport({
  startDate: '2026-03-01',
  endDate: '2026-03-31',
  groupBy: 'tag',
  tags: ['team:finance'],
});
 
for (const row of report.results) {
  console.log(
    `${row.tag}: $${row.totalCost.toFixed(4)} (${row.requestCount} requests)`,
  );
}

See the 创意脚本 docs on spend reports for the full list of parameters and response fields.

Use the 创意脚本's getGenerationInfo() method to look up a specific generation by its ID, including cost, token usage, latency, and provider details. For the dedicated workflow and REST API links, see Generation Lookup. Generation IDs are available in providerMetadata.gateway.generationId on both generateText and streamText responses.

When streaming, the generation ID is injected on the first content chunk, so you can capture it early without waiting for completion. This is useful when a network interruption cuts off the final response. 资产生成 records the final status server-side, so you can use the generation ID to look up the results later.

import { gateway, generateText } from 'ai';
 
const result = await generateText({
  model: 'anthropic/claude-opus-5',
  prompt: 'Explain quantum entanglement briefly',
});
 
const generationId = result.providerMetadata?.gateway?.generationId;
if (typeof generationId !== 'string') throw new Error('Missing generation ID');
const generation = await gateway.getGenerationInfo({ id: generationId });
 
console.log(`Model: ${generation.model}`);
console.log(`Cost: $${generation.totalCost.toFixed(6)}`);
console.log(`Latency: ${generation.latency}ms`);
console.log(`Prompt tokens: ${generation.promptTokens}`);
console.log(`Completion tokens: ${generation.completionTokens}`);
import { gateway, streamText } from 'ai';
 
const result = streamText({
  model: 'anthropic/claude-opus-5',
  prompt: 'Explain quantum entanglement briefly',
});
 
let generationId: string | undefined;
 
for await (const part of result.stream) {
  if (
    !generationId &&
    'providerMetadata' in part &&
    typeof part.providerMetadata?.gateway?.generationId === 'string'
  ) {
    generationId = part.providerMetadata.gateway.generationId;
  }
}
 
if (generationId) {
  const generation = await gateway.getGenerationInfo({ id: generationId });
  console.log(`Cost: $${generation.totalCost.toFixed(6)}`);
  console.log(`Finish reason: ${generation.finishReason}`);
}

See the 创意脚本 docs on generation lookup for the full list of response fields.

curl "https://ai-gateway.vercel.sh/v1/report?start_date=2026-01-01&end_date=2026-01-31&date_part=day" \
  -H "Authorization: Bearer YOUR_API_KEY"
curl "https://ai-gateway.vercel.sh/v1/report?start_date=2026-01-01&end_date=2026-01-31&date_part=hour&group_by=model" \
  -H "Authorization: Bearer YOUR_API_KEY"
curl "https://ai-gateway.vercel.sh/v1/report?start_date=2026-01-01&end_date=2026-01-31&group_by=user" \
  -H "Authorization: Bearer YOUR_API_KEY"
curl "https://ai-gateway.vercel.sh/v1/report?start_date=2026-01-01&end_date=2026-01-31&group_by=tag" \
  -H "Authorization: Bearer YOUR_API_KEY"
curl "https://ai-gateway.vercel.sh/v1/report?start_date=2026-01-01&end_date=2026-01-31&group_by=credential_type" \
  -H "Authorization: Bearer YOUR_API_KEY"

When you authenticate the report with an 资产生成 API key, use self to return only spend attributed to that key:

curl "https://ai-gateway.vercel.sh/v1/report?start_date=2026-01-01&end_date=2026-01-31&api_key_id=self" \
  -H "Authorization: Bearer $AI_GATEWAY_API_KEY"

You can combine filters to narrow results:

curl "https://ai-gateway.vercel.sh/v1/report?start_date=2026-01-01&end_date=2026-01-31&date_part=day&user_id=user_123&model=anthropic/claude-sonnet-5&tags=production,api" \
  -H "Authorization: Bearer YOUR_API_KEY"

Attach observability tags to a virtual model to label all of its traffic. The virtual model's tags merge with the request's tags instead of replacing them. Requests cap at ten tags, and a virtual model's tags merge on top without re-checking the cap.

Last updated September 11, 2026

Was this helpful?

supported.