|
This version is still in development and is not considered stable yet. For the latest stable version, please use Spring AI 2.0.1! |
Chat Response Metadata
Use of an AI, such as OpenAI’s ChatGPT, consumes resources and generates metrics returned by the AI provider based on the usage and requests made to the AI through the API. Consumption is typically in the form of requests made or tokens used in a given timeframe, such as monthly, that AI providers use to measure this consumption and reset limits. Your rate limits are directly determined by your plan when you signed up with your AI provider. For instance, you can review details on OpenAI’s rate limits and plans by following the links.
To help garner insight into your AI (model) consumption and general usage, Spring AI provides an API to introspect the metadata that is returned by AI providers in their APIs.
Spring AI defines 3 primary types to examine these metrics: ChatResponseMetadata, RateLimit and Usage.
All of these types can be accessed programmatically from the ChatResponse returned and initiated by an AI request.
ChatResponseMetadata
An instance of ChatResponseMetadata is automatically created by Spring AI when an AI request is made through the AI provider’s API and an AI response is returned.
You can get access to the AI provider metadata from the ChatResponse using:
ChatResponseMetadata from ChatResponseChatResponse response = chatModel.call(prompt);
ChatResponseMetadata metadata = response.getMetadata();
String id = metadata.getId();
String model = metadata.getModel();
RateLimit rateLimit = metadata.getRateLimit();
Usage usage = metadata.getUsage();
PromptMetadata promptMetadata = metadata.getPromptMetadata();
You might imagine that you can rate limit your own Spring applications using AI, or restrict Prompt sizes, which affect your token usage, in an automated, intelligent and realtime manner.
Minimally, you can simply gather these metrics to monitor and report on your consumption.
Rate Limit
The RateLimit interface provides access to actual information returned by an AI provider on your API usage when making AI requests.
RateLimit interfaceinterface RateLimit {
Long getRequestsLimit();
Long getRequestsRemaining();
Duration getRequestsReset();
Long getTokensLimit();
Long getTokensRemaining();
Duration getTokensReset();
}
requestsLimit and requestsRemaining let you know how many AI requests, based on the AI provider plan you chose when you signed up, that you can make in total along with your remaining balance within the given timeframe.
requestsReset returns a Duration of time before the timeframe expires and your limits reset based on your chosen plan.
The methods for tokensLimit, tokensRemaining and tokensReset are similar, but focus on token limits, balance and resets instead.
RateLimit from ChatResponseMetadataRateLimit rateLimit = metadata.getRateLimit();
Long tokensRemaining = rateLimit.getTokensRemaining();
// do something interesting with the RateLimit metadata
AI providers such as OpenAI return rate limit information in the HTTP headers of their REST API responses.
Where a model implementation maps those headers, they are exposed through RateLimit.
Support varies by model implementation.
When rate limit information is not available, getRateLimit() returns an EmptyRateLimit whose accessors return zero rather than null.
Use instanceof EmptyRateLimit to distinguish an absent rate limit from a real zero value.
Usage
Usage data can be obtained from the ChatResponseMetadata object.
The Usage interface is defined as:
Usage interfaceinterface Usage {
Integer getPromptTokens();
Integer getCompletionTokens();
default Integer getTotalTokens() { ... }
Object getNativeUsage();
}
The method names are self-explanatory: they tell you the tokens that the AI required to process the Prompt and generate a response.
totalTokens is the sum of promptTokens and completionTokens.
Spring AI computes this by default, but the information is also returned in the AI response from providers like OpenAI.
Using with ChatModel
Here’s a complete example showing how to track usage with OpenAI’s ChatModel:
@SpringBootConfiguration
public class Configuration {
@Bean
public OpenAiChatModel openAiChatModel() {
return OpenAiChatModel.builder()
.options(OpenAiChatOptions.builder()
.apiKey(System.getenv("OPENAI_API_KEY"))
.build())
.build();
}
}
@Service
public class ChatService {
private final OpenAiChatModel chatModel;
public ChatService(OpenAiChatModel chatModel) {
this.chatModel = chatModel;
}
public void demonstrateUsage() {
// Create a chat prompt
Prompt prompt = new Prompt("What is the weather like today?");
// Get the chat response
ChatResponse response = this.chatModel.call(prompt);
// Access the usage information
Usage usage = response.getMetadata().getUsage();
// Get standard usage metrics
System.out.println("Prompt Tokens: " + usage.getPromptTokens());
System.out.println("Completion Tokens: " + usage.getCompletionTokens());
System.out.println("Total Tokens: " + usage.getTotalTokens());
// Access native OpenAI usage data with detailed token information
if (usage.getNativeUsage() instanceof com.openai.models.completions.CompletionUsage nativeUsage) {
// Detailed prompt token information
nativeUsage.promptTokensDetails().ifPresent(details -> {
System.out.println("Prompt Tokens Details:");
details.audioTokens().ifPresent(tokens -> System.out.println("- Audio Tokens: " + tokens));
details.cachedTokens().ifPresent(tokens -> System.out.println("- Cached Tokens: " + tokens));
});
// Detailed completion token information
nativeUsage.completionTokensDetails().ifPresent(details -> {
System.out.println("Completion Tokens Details:");
details.reasoningTokens().ifPresent(tokens -> System.out.println("- Reasoning Tokens: " + tokens));
details.acceptedPredictionTokens().ifPresent(tokens -> System.out.println("- Accepted Prediction Tokens: " + tokens));
details.audioTokens().ifPresent(tokens -> System.out.println("- Audio Tokens: " + tokens));
details.rejectedPredictionTokens().ifPresent(tokens -> System.out.println("- Rejected Prediction Tokens: " + tokens));
});
}
}
}
Prompt Cache Usage Metrics
For providers that support prompt caching, the Usage interface provides unified access to cache metrics without requiring provider-specific casting:
Usage usage = response.getMetadata().getUsage();
// Unified cache metrics — works across all providers
Long cacheReadTokens = usage.getCacheReadInputTokens();
Long cacheWriteTokens = usage.getCacheWriteInputTokens();
if (cacheReadTokens != null && cacheReadTokens > 0) {
System.out.println("Cache hit: " + cacheReadTokens + " tokens read from cache");
}
if (cacheWriteTokens != null && cacheWriteTokens > 0) {
System.out.println("Cache write: " + cacheWriteTokens + " tokens written to cache");
}
These methods return null for providers that do not support prompt caching.
The following table shows prompt cache metrics availability by provider:
| Provider | Cache Read Tokens | Cache Write Tokens |
|---|---|---|
Anthropic |
Yes |
Yes ( |
AWS Bedrock |
Yes |
Yes |
OpenAI |
Yes ( |
No |
Google Gemini |
Yes ( |
No |
DeepSeek |
No |
No |
Mistral |
No |
No |
Ollama |
No |
No |
For detailed provider-specific cache metrics (such as per-modality cache breakdowns in Gemini), use getNativeUsage() to access the provider’s native usage object.
|
Cumulative Usage Across Multi-Step Flows
When a response is produced through a multi-step flow such as a tool-calling loop, getUsage() reports the cumulative token usage across every model call in that exchange, not just the last call.
For example, a ChatClient conversation that triggers one tool call performs at least two model calls, and the returned getUsage() reflects the sum of both.
ChatResponse response = chatClient.prompt("What is the weather in Paris?")
.tools(new WeatherTools())
.call()
.chatResponse();
// Cumulative across all model calls in the tool-calling loop
Usage usage = response.getMetadata().getUsage();
int totalTokens = usage.getTotalTokens();
The cumulative total is computed via org.springframework.ai.support.UsageCalculator, which sums the standard token counts and the unified cache metrics (getCacheReadInputTokens() / getCacheWriteInputTokens()).
Provider-specific native usage objects cannot be merged across responses.
As a result, getNativeUsage() returns null once usage has been aggregated across more than one model call (for example after a tool-calling loop); it is preserved only for single-call responses.
If you need the provider’s native usage object, read it from an individual ChatModel call rather than from a multi-step ChatClient exchange.
|