This version is still in development and is not considered stable yet. For the latest stable version, please use Spring AI 2.0.1!

Chat Response Metadata

Use of an AI, such as OpenAI’s ChatGPT, consumes resources and generates metrics returned by the AI provider based on the usage and requests made to the AI through the API. Consumption is typically in the form of requests made or tokens used in a given timeframe, such as monthly, that AI providers use to measure this consumption and reset limits. Your rate limits are directly determined by your plan when you signed up with your AI provider. For instance, you can review details on OpenAI’s rate limits and plans by following the links.

To help garner insight into your AI (model) consumption and general usage, Spring AI provides an API to introspect the metadata that is returned by AI providers in their APIs.

Spring AI defines 3 primary types to examine these metrics: ChatResponseMetadata, RateLimit and Usage. All of these types can be accessed programmatically from the ChatResponse returned and initiated by an AI request.

ChatResponseMetadata

An instance of ChatResponseMetadata is automatically created by Spring AI when an AI request is made through the AI provider’s API and an AI response is returned. You can get access to the AI provider metadata from the ChatResponse using:

Get access to ChatResponseMetadata from ChatResponse
ChatResponse response = chatModel.call(prompt);

ChatResponseMetadata metadata = response.getMetadata();

String id = metadata.getId();
String model = metadata.getModel();
RateLimit rateLimit = metadata.getRateLimit();
Usage usage = metadata.getUsage();
PromptMetadata promptMetadata = metadata.getPromptMetadata();

You might imagine that you can rate limit your own Spring applications using AI, or restrict Prompt sizes, which affect your token usage, in an automated, intelligent and realtime manner. Minimally, you can simply gather these metrics to monitor and report on your consumption.

Rate Limit

The RateLimit interface provides access to actual information returned by an AI provider on your API usage when making AI requests.

RateLimit interface
interface RateLimit {

	Long getRequestsLimit();

	Long getRequestsRemaining();

	Duration getRequestsReset();

	Long getTokensLimit();

	Long getTokensRemaining();

	Duration getTokensReset();

}

requestsLimit and requestsRemaining let you know how many AI requests, based on the AI provider plan you chose when you signed up, that you can make in total along with your remaining balance within the given timeframe. requestsReset returns a Duration of time before the timeframe expires and your limits reset based on your chosen plan. The methods for tokensLimit, tokensRemaining and tokensReset are similar, but focus on token limits, balance and resets instead.

Get access to RateLimit from ChatResponseMetadata
RateLimit rateLimit = metadata.getRateLimit();

Long tokensRemaining = rateLimit.getTokensRemaining();

// do something interesting with the RateLimit metadata

AI providers such as OpenAI return rate limit information in the HTTP headers of their REST API responses. Where a model implementation maps those headers, they are exposed through RateLimit.

Support varies by model implementation. When rate limit information is not available, getRateLimit() returns an EmptyRateLimit whose accessors return zero rather than null. Use instanceof EmptyRateLimit to distinguish an absent rate limit from a real zero value.

Usage

Usage data can be obtained from the ChatResponseMetadata object. The Usage interface is defined as:

Usage interface
interface Usage {

	Integer getPromptTokens();

	Integer getCompletionTokens();

	default Integer getTotalTokens() { ... }

	Object getNativeUsage();

}

The method names are self-explanatory: they tell you the tokens that the AI required to process the Prompt and generate a response. totalTokens is the sum of promptTokens and completionTokens. Spring AI computes this by default, but the information is also returned in the AI response from providers like OpenAI.

Using with ChatModel

Here’s a complete example showing how to track usage with OpenAI’s ChatModel:

@SpringBootConfiguration
public class Configuration {

	@Bean
	public OpenAiChatModel openAiChatModel() {
		return OpenAiChatModel.builder()
			.options(OpenAiChatOptions.builder()
				.apiKey(System.getenv("OPENAI_API_KEY"))
				.build())
			.build();
	}

}

@Service
public class ChatService {

	private final OpenAiChatModel chatModel;

	public ChatService(OpenAiChatModel chatModel) {
		this.chatModel = chatModel;
	}

	public void demonstrateUsage() {
		// Create a chat prompt
		Prompt prompt = new Prompt("What is the weather like today?");

		// Get the chat response
		ChatResponse response = this.chatModel.call(prompt);

		// Access the usage information
		Usage usage = response.getMetadata().getUsage();

		// Get standard usage metrics
		System.out.println("Prompt Tokens: " + usage.getPromptTokens());
		System.out.println("Completion Tokens: " + usage.getCompletionTokens());
		System.out.println("Total Tokens: " + usage.getTotalTokens());

		// Access native OpenAI usage data with detailed token information
		if (usage.getNativeUsage() instanceof com.openai.models.completions.CompletionUsage nativeUsage) {

			// Detailed prompt token information
			nativeUsage.promptTokensDetails().ifPresent(details -> {
				System.out.println("Prompt Tokens Details:");
				details.audioTokens().ifPresent(tokens -> System.out.println("- Audio Tokens: " + tokens));
				details.cachedTokens().ifPresent(tokens -> System.out.println("- Cached Tokens: " + tokens));
			});

			// Detailed completion token information
			nativeUsage.completionTokensDetails().ifPresent(details -> {
				System.out.println("Completion Tokens Details:");
				details.reasoningTokens().ifPresent(tokens -> System.out.println("- Reasoning Tokens: " + tokens));
				details.acceptedPredictionTokens().ifPresent(tokens -> System.out.println("- Accepted Prediction Tokens: " + tokens));
				details.audioTokens().ifPresent(tokens -> System.out.println("- Audio Tokens: " + tokens));
				details.rejectedPredictionTokens().ifPresent(tokens -> System.out.println("- Rejected Prediction Tokens: " + tokens));
			});
		}
	}
}

Using with ChatClient

If you are using the ChatClient, you can access the usage information using the ChatResponse object:

ChatResponse response = chatClient.prompt("What is the weather like today?")
		.call()
		.chatResponse();

Usage usage = response.getMetadata().getUsage();

Prompt Cache Usage Metrics

For providers that support prompt caching, the Usage interface provides unified access to cache metrics without requiring provider-specific casting:

Usage usage = response.getMetadata().getUsage();

// Unified cache metrics — works across all providers
Long cacheReadTokens = usage.getCacheReadInputTokens();
Long cacheWriteTokens = usage.getCacheWriteInputTokens();

if (cacheReadTokens != null && cacheReadTokens > 0) {
	System.out.println("Cache hit: " + cacheReadTokens + " tokens read from cache");
}
if (cacheWriteTokens != null && cacheWriteTokens > 0) {
	System.out.println("Cache write: " + cacheWriteTokens + " tokens written to cache");
}

These methods return null for providers that do not support prompt caching.

The following table shows prompt cache metrics availability by provider:

Provider Cache Read Tokens Cache Write Tokens

Anthropic

Yes

Yes (cacheCreationInputTokens)

AWS Bedrock

Yes

Yes

OpenAI

Yes (cachedTokens)

No

Google Gemini

Yes (cachedContentTokenCount)

No

DeepSeek

No

No

Mistral

No

No

Ollama

No

No

For detailed provider-specific cache metrics (such as per-modality cache breakdowns in Gemini), use getNativeUsage() to access the provider’s native usage object.

Cumulative Usage Across Multi-Step Flows

When a response is produced through a multi-step flow such as a tool-calling loop, getUsage() reports the cumulative token usage across every model call in that exchange, not just the last call. For example, a ChatClient conversation that triggers one tool call performs at least two model calls, and the returned getUsage() reflects the sum of both.

ChatResponse response = chatClient.prompt("What is the weather in Paris?")
		.tools(new WeatherTools())
		.call()
		.chatResponse();

// Cumulative across all model calls in the tool-calling loop
Usage usage = response.getMetadata().getUsage();
int totalTokens = usage.getTotalTokens();

The cumulative total is computed via org.springframework.ai.support.UsageCalculator, which sums the standard token counts and the unified cache metrics (getCacheReadInputTokens() / getCacheWriteInputTokens()).

Provider-specific native usage objects cannot be merged across responses. As a result, getNativeUsage() returns null once usage has been aggregated across more than one model call (for example after a tool-calling loop); it is preserved only for single-call responses. If you need the provider’s native usage object, read it from an individual ChatModel call rather than from a multi-step ChatClient exchange.