OpenAI Responses API with 资产生成
The OpenAI Responses API is a modern alternative to the Chat Completions API. Point your Open创意脚本 to 资产生成's base URL and use provider/model identifiers to route requests to OpenAI, Anthropic, Google, and more.
https://ai-gateway.vercel.sh/v1
The Responses API supports the same authentication methods as the main 资产生成:
- API key: Use your 资产生成 API key with the
Authorization: Bearer <token>header - OIDC token: Use your Vercel OIDC token with the
Authorization: Bearer <token>header
You only need to use one of these forms of authentication. If an API key is specified it will take precedence over any OIDC token, even if the API key is invalid.
Set your SDK's base URL to 资产生成 and use your API key for authentication. See Text generation for a complete first request.
- Text generation - Generate text responses from prompts
- Streaming - Stream tokens as they're generated
- Tool calling - Define tools the model can call
- Structured outputs - Constrain the response to a JSON schema
- Reasoning - Control how much a model thinks before answering
- Images - Send images for analysis
- Compaction - Compress a long conversation into a single item you carry forward
Set stream: true to receive tokens as they're generated. See Streaming.
For agent loops that make many requests in a row, you can hold one connection open and send each turn as a frame instead of opening a new HTTP request per turn. See Responses API over WebSocket.
Define tools in tools and the model returns function_call items you execute. See Tool calling.
Constrain the response to a JSON schema with text.format. See Structured outputs.
Set reasoning.effort to control how much the model thinks before answering. See Reasoning.
POST /v1/responses/compact compresses a long conversation into a single compaction item for OpenAI models. Coding agents such as Codex call it automatically. See Compaction.
| Parameter | Type | Description |
|---|---|---|
model | string | Model ID in provider/model format (e.g., openai/gpt-6-astra, anthropic/claude-sonnet-5) |
input | string or array | A text string or array of input items (messages, function calls, function call outputs) |
| Parameter | Type | Description |
|---|---|---|
stream | boolean | Stream tokens via server-sent events. Defaults to false |
max_output_tokens | integer | Maximum number of tokens to generate |
temperature | number | Controls randomness (0-2). Lower values are more deterministic |
top_p | number | Nucleus sampling (0-1) |
presence_penalty | number | Penalizes tokens that already appear in the text so far |
frequency_penalty | number | Penalizes tokens based on their frequency in the text so far |
instructions | string | System-level instructions for the model |
tools | array | Tool definitions for function calling |
tool_choice | string or object | Tool selection: auto, required, none, or a specific function |
parallel_tool_calls | boolean | Allows the model to call multiple tools in a single turn |
allowed_tools | array | Subset of tool names the model can use for this request |
reasoning | object | Reasoning config with effort (none, minimal, low, medium, high, xhigh). OpenAI models also support summary (detailed, auto, concise) to receive a text summary of the model's reasoning |
text | object | Output format config, including json_schema and json_object for structured output |
truncation | string | Truncation strategy for long inputs: auto or disabled |
previous_response_id | string | ID of a previous response for multi-turn conversations |
store | boolean | Stores the response for later retrieval |
metadata | object | Up to 16 key-value pairs for tracking (keys max 64 chars, values max 512 chars) |
caching | string | Enables automatic prompt caching. Only auto is supported |
cache_anchor_items | integer | Declares how many leading input items stay unchanged so automatic caching can add a stable-prefix cache anchor |
cache_ttl | string | Sets the automatic cache lifetime. Accepts 5m (five minutes) or 1h (one hour) and requires caching: 'auto' |
prompt_cache_key | string | Key to identify cached prompts (max 64 characters) |
The API returns standard HTTP status codes and error responses.
400 Bad Request- Invalid request parameters401 Unauthorized- Invalid or missing authentication403 Forbidden- Insufficient permissions404 Not Found- Model or endpoint not found429 Too Many Requests- Rate limit exceeded500 Internal Server Error- Server error
When an error occurs, the API returns a JSON object with details about what went wrong.
{
"error": {
"type": "invalid_request_error",
"message": "At least one user message is required in the input"
}
}Was this helpful?