Skip to content

Anthropic starts charging for some of Claude's safety refusals

Anthropic now bills Claude API requests refused by its safety classifiers before any output when tagged bio, frontier_llm, or reasoning_extraction, aiming to make large-scale safeguard probing expensive.

By VibecodedThis 2 min read
Anthropic wordmark logo on a white background
Anthropic, via Wikimedia Commons

Anthropic is changing what a "no" costs. The company has begun charging developers for Claude API requests that its safety classifiers refuse before the model produces any output, when the refusal falls into one of three classifier categories. The change is documented in Anthropic's refusals and fallback documentation, which lays out exactly which refusals are billed and which are not.

What is billed

A refusal that arrives before any output is now billed when the classifier tags it as bio (requests that could enable biological harm, like dangerous lab methods), frontier_llm (requests that could assist development of competing AI models), or reasoning_extraction (requests asking the model to reproduce its internal reasoning as output text). These are the categories where Anthropic says it measures low false positive rates as of September 2026.

Refusals tagged cyber or general_harms, and refusals with no category at all, remain free. Either way the request still counts against your rate limits, and the response comes back as a normal HTTP 200 with stop_reason: "refusal", an empty content array, and token counts in usage. You are paying for a call that returned no text.

Anthropic's stated rationale is blunt: billing pre-output refusals is meant "to disrupt attempts to circumvent Anthropic's safeguards at scale." The company attributes the change to coordinated probing attacks observed in recent weeks, according to AlphaSignal's writeup of the announcement. The logic is that automated jailbreak campaigns that hammer the API thousands of times should get expensive fast.

The false positive problem

The wrinkle is that classifiers are not perfect, and Anthropic's own documentation admits benign work can trip these labels. Beneficial life sciences work can trigger bio. Benign machine learning work can trigger frontier_llm. Both now carry a charge when they do. As MIXED News notes, that means a legitimate request can cost you money because a classifier misread it, and the stop_details.category in the response is the clearest record you get of why you were charged.

A few billing mechanics worth knowing from the docs:

  • Mid-stream refusals, which arrive after Claude has started writing, were already billed at normal rates for input tokens plus whatever output streamed. That is unchanged.
  • If you use Anthropic's fallback feature to re-serve a refused request on another model, the refusal is billed in addition to the fallback request when it arrived mid-stream or falls in a billed category.
  • The rules apply everywhere: the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry.
  • The billed category list can change as Anthropic keeps measuring false positive rates.

What developers should do

Start logging stop_details.category on refusals alongside your usage records, so you can tell a charged refusal from a free one and spot false positive patterns in your own workloads. If you run life sciences or ML research pipelines through Claude, budget for the possibility that some fraction of legitimate requests now costs money, and consider the fallback parameter for requests that get refused mid-stream rather than retrying blindly. The era of the free refusal is ending, at least for the three categories Anthropic watches most closely.