
A listed model price does not describe the full cost of helping a website visitor. Vultr’s September 15, 2026 announcement about cloud and AI economics provides a useful starting point: measure an outcome your service actually delivers, such as an accepted support answer.
1. Define the outcome and boundary
Select one use case, for example answering questions from approved documentation. Define an accepted answer with observable criteria: correct information, the question addressed and no correction needed under your review rubric. A successful HTTP request is not by itself an accepted answer.
Choose the same reporting period and workload boundary for both spending and outcomes. Separate internal testing from visitor traffic. If quality is assessed through a sample, record its size and selection method instead of presenting the result as an exhaustive review.
Give each request a technical identifier without embedding personal information. Associate subsequent model calls, retrieval operations and retries with that identifier. This makes it possible to see why one apparently simple visitor question required several paid operations.
2. Reconcile usage with observed charges
Collect actual counters and charges exposed by your providers. Keep model usage, retrieval services and the allocated hosting share separate. Retain the original currency; if you convert it for reporting, record the exchange rate and the date used.
A tracking worksheet might use these fields:
request_id,model_calls,retries,model_costretrieval_cost,hosting_share,acceptedThis is a suggested record structure, not a provider-specific export format. Billing units differ between services. Use the units shown in your own statement rather than asking an assistant to invent an estimated rate.
Avoid counting shared infrastructure twice. When the same server also hosts the website, select and document an allocation method. Label the allocated amount as an estimate if it is not directly measured. List excluded items, such as human review time, so someone comparing results understands the boundary.
3. Compare equivalent workloads
Divide the total cost within your chosen boundary by accepted answers in the same period. Keep unsuccessful attempts associated with that use case in the numerator. If no answer was accepted, the ratio is undefined; reporting zero would suggest the opposite of what happened.
Compare two configurations using the same questions and acceptance rules. Review answer quality and response time alongside cost. A model with a lower per-call price may require extra attempts, and the measurement should make those attempts visible.
Keep exclusions and estimates attached to every reported result. The result is a repeatable indicator for your own assistant, not a promise based on a catalogue price. As traffic and document quality change, repeat the measurement using the same definitions so that differences in the figures reflect meaningful operational changes.
Sources: Vultr