New SAP AI Core Calculator 2.0
Boidra Admin9 min read
SAP AI Core Cost Calculator 2.0: Understanding AI Agent Costs
SAP AI Core costs include more than model calls. This article explains how tokens become Capacity Units, breaks down a monthly estimate, and compares model consumption. It also explains prompt optimization and why translation, testing, and other services belong in the budget.
The SAP AI Core Cost Calculator estimates monthly use in Capacity Units, or CUs. Version 2.0 adds nine starting templates, an expanded model catalog, and PDF export.
The calculator showed 68 catalog entries when checked on September 28, 2026. These include different versions, context limits, and model types. They are not 68 interchangeable models for agents.
All calculator figures below are estimates from that date.
How tokens become CUs
SAP separates three units:
| Unit | Meaning |
|---|---|
| Model tokens | Pieces of input or output processed by a model |
| GenAI tokens | SAP's common consumption unit across models |
| Capacity Units (CUs) | The unit used for SAP AI Core billing |
Input and output have separate conversion rates. The rates depend on the model. SAP points to SAP Note 3437766 for model details and token conversion rates.
The calculation has two steps:
GenAI tokens per request =
(input tokens / 1,000 × input conversion rate)
+ (output tokens / 1,000 × output conversion rate)
CUs per request = GenAI tokens per request × CU conversion factorSAP's Metering and Pricing for Generative AI uses 1.90385 as the CU conversion factor in its example. This is a consumption factor, not a euro or dollar price.
The page labels the example values as fictitious. Use it to understand the calculation, and check the current rates for your selected model before budgeting.
A worked conversion example
Using the values in SAP's example:
| Input | Value |
|---|---|
| Input tokens per request | 3,500 |
| Output tokens per request | 300 |
| Input conversion rate per 1,000 tokens | 0.00112 |
| Output conversion rate per 1,000 tokens | 0.00320 |
| CU conversion factor | 1.90385 |
| Requests per month | 25,000 |
Input: 3,500 / 1,000 × 0.00112 = 0.00392 GenAI tokens
Output: 300 / 1,000 × 0.00320 = 0.00096 GenAI tokens
Total per request = 0.00488 GenAI tokens
CUs per request = 0.00488 × 1.90385 = 0.009290788
Monthly CUs = 25,000 × 0.009290788 = 232.2697The result is about 232.27 CU per month for model calls, before supporting services.
SAP shows 232.25 because it rounds each request to 0.00929 CU first. Its calculation also prints a dollar sign at that step, although the formula calculates CUs. A currency price must be applied separately. Source: SAP's worked example.
How much does one CU cost?
The calculator states that your SAP contract defines the currency rate.
Monthly cost = monthly CUs × agreed price per CUFor example, at an assumed price of €1 per CU, 2,112.36 CU would cost €2,112.36. This is an illustration, not a confirmed price for your account.
Check your agreement and the SAP Discovery Center Estimator for the applicable price. Do not apply 1.90385 again to totals that are already shown in CUs.
SAP AI Core CUs are also different from the AI Units used for SAP Business AI offerings. See SAP Business AI pricing.
A monthly cost breakdown
The example below shows why model pricing alone is not enough.
| Component | CU/month | Share of known subtotal |
|---|---|---|
| Foundation Models | 331.41 | 15.69% |
| Orchestration | 817.34 | 38.69% |
| Prompt Optimization | 958.15 | 45.36% |
| Evaluations | 5.46 | 0.26% |
| Compute Instances | 0.00 | 0.00% |
| Object Storage | Not supplied | Not included |
| Known subtotal | 2,112.36 | 100% |
If Object Storage is zero and there are no other charges, the total is 2,112.36 CU per month.
Model calls account for 15.7% of this estimate. Orchestration and prompt optimization account for 84.1%. This split comes from the selected settings. It is not a standard split for every SAP AI application.
What is included in orchestration?
Orchestration provides a common API for model access and supporting services. These services can retrieve business data, filter content, translate text, mask personal data, and record requests.
The following settings reproduced the example's 817.34 CU orchestration total in the calculator. They are one matching configuration, not proof of the original settings.
| Service | Tested monthly settings | Displayed CU/month |
|---|---|---|
| Content filtering | 10,000 text blocks | 5.71 |
| Grounding | 10,000 queries; storage field set to 1 GB/day | 51.28 |
| Translation | 10,000 requests | 740.22 |
| Data masking | 5,000 requests | 11.23 |
| Inference observability | 10,000 requests | 8.90 |
| Total | 817.34 |
Translation makes up 90.6% of this orchestration estimate. The calculator assumes about 24 text blocks per translation request. Set translation volume to match how often your application will actually use it.
The conversion factor helps explain several displayed totals:
Filtering:
10,000 × 0.00030 × 1.90385 = 5.71155 CU
Translation:
10,000 × 24 × 0.00162 × 1.90385 = 740.21688 CU
Data masking:
5,000 × 0.00118 × 1.90385 = 11.232715 CUThese calculations match the rounded calculator results. The calculator's helper text labels the starting values as CU rates, but the displayed totals behave as though a further 1.90385 conversion is applied. This is an inference from the tested numbers, not a confirmed definition of those labels.
Do not apply that multiplier to every service. Storage and observability have their own measures and conversion rules.
The grounding storage amount still needs care. The tested total is consistent with about 51.07 CU for retrieval plus 0.21 CU for storage. It does not appear to multiply the storage field by 30 days. SAP's documentation treats grounding storage as gigabyte-days. Check the storage period when preparing a monthly estimate. See SAP's metering guide.
What is prompt optimization in SAP AI Core?
Prompt optimization is a service that improves the instructions you send to a model. You provide a starting prompt, example data, a target model, and a way to score the answers. SAP AI Core runs an optimization job that tests and refines the prompt against that score.
It changes the prompt rather than retraining the model.
A typical process is:
- Prepare a prompt template and representative examples with expected answers.
- Choose the model you want to use.
- Choose a metric that measures whether the answers are useful.
- Run the optimization job.
- Review the resulting prompt and test it on separate examples before using it.
The optimized prompt can be saved and reused. SAP provides the workflow through SAP AI Core and SAP AI Launchpad. See SAP's prompt optimization tutorial.
For example, an agent might read a support request and return its category, urgency, and responsible team. Prompt optimization can help improve those outputs against a set of correctly labelled requests. It can also be useful when moving an existing prompt to another model.
For agents that call tools, SAP provides a tool-calling optimization tutorial. It shows how to compare a starting prompt with an optimized prompt using test data and a metric for structured output.
An improved test score does not guarantee lower costs or better results on every request. Check answer quality, prompt length, and running cost before adopting the new prompt.
How prompt optimization and evaluations affect the estimate
The calculator reproduced the example values with these inputs:
| Component | Monthly input | Displayed CU/month | Rate derived from this test |
|---|---|---|---|
| Prompt Optimization | 25 optimizations | 958.15 | 38.326 CU per optimization |
| Evaluations | 10 node-hours | 5.46 | 0.546 CU per node-hour |
These rates are derived from the calculator settings. They are not a separate confirmation of contractual billing terms. Source: SAP AI Core Cost Calculator.
The 25 optimizations are optimization jobs, not 25 user questions. Budget for how often you plan to improve prompts. A month spent developing or migrating prompts may need more optimization work than a month of normal operation.
Evaluations measure how well a prompt and model perform. Prompt optimization uses scoring to improve the prompt; evaluations let you check and compare results. SAP supports separate training and test datasets for optimization. See SAP AI Core release notes.
The calculator estimates evaluation use in node-hours. That is why Evaluations can have a cost even when the separate Compute Instances line is zero.
For managed foundation models, zero custom compute does not mean free model calls. Token charges still apply. Custom model workloads can also have compute, storage, and baseline charges. See SAP AI Core metering.
Comparing models for agents
The catalog offers a popularity sort, but it does not provide enough evidence here to rank models by actual agent usage. The table below is a practical comparison of selected models.
Each model was tested in the calculator with:
- 1,000 requests per month
- 1,000 input tokens per request
- 1,000 output tokens per request
- Standard pricing mode
This is one million input tokens plus one million output tokens per month. The figures include only the model line item.
| Model | CU/month | Possible use to test |
|---|---|---|
| GPT-5.4 Nano | 1.75 | Classification, routing, and simple extraction |
| GPT-5 Mini | 2.95 | Routine agent steps and structured tasks |
| Gemini 3.5 Flash Lite | 6.04 | High-volume tasks with text or other inputs |
| Claude 4.5 Haiku | 8.49 | Short assistant tasks |
| Gemini 3.5 Flash | 21.89 | Document and multimodal tasks |
| Claude 4.6 Sonnet | 24.94 | General agent tasks and coding |
| Claude 4.8 Opus | 41.37 | Difficult tasks that need more reasoning |
| GPT 5.6 Sol, up to 272k context tier | 44.87 | Complex planning and agent tasks |
Source: SAP calculator, checked September 28, 2026. The suggested uses are starting points for testing, not performance results.
These are not prices for one million input tokens alone. Different input and output volumes, caching, and context tiers can change the comparison. Model availability also depends on the region.
Other catalog entries do different jobs. Embedding models help find relevant content. Cohere Reranker helps order search results. SAP RPT and Prior Labs TabPFN handle tabular tasks. They can support an agent without replacing its main conversational model.
Budget for the whole task
An agent may call a model several times to finish one task. It may plan an action, call a tool, read the result, retry, and write an answer. Each model call adds tokens.
Monthly model calls = tasks per month × average model calls per taskFor example, 10,000 tasks with five model calls each produce 50,000 calls before extra retries. Each step can use a different model and a different number of tokens.
Start by testing a smaller model for routine steps. Use a more capable model where it improves results enough to justify the cost. Compare cost per completed task as well as accuracy and response time.
In this example, translation and prompt optimization together account for 1,698.37 CU per month, or 80.4% of the known subtotal. Review those two settings first. They have more effect on this estimate than a small change in model price.
Keep reading
Telemetry in AI: why it matters and how it improves testing
You cannot test what you cannot see. AI telemetry - a span for every agent step - turns non-deterministic systems into something you can measure, regression-test and trust.
Joule Studio Classic vs the new Joule Studio: what changed
Joule Studio went from low-code skill building (2024) to autonomous agent building (2025) to intent-based development in Joule Studio 2.0 (2026). Here is the progression and what it means for you.
How SAP Joule works: from copilot to a network of agents
Joule is SAP’s AI copilot - natively embedded in SAP apps and grounded in business data. In 2025–2026 it grew from a single assistant into an orchestrated network of specialized agents.