Skip to main content
Contact Sales (or support@vantage.sh/your account team) to enable LLM Enrichment for your account and discuss availability.
Vantage reads per-request token usage from your LiteLLM proxy and joins it to your model-provider costs. A Vantage callback loaded inside your LiteLLM process records one usage event per request (the upstream provider, model, token counts, and the metadata you attach) and hands it to the Vantage collector, a sidecar that batches and uploads that usage to Vantage. Vantage splits each matching cost row into enriched rows by token share and adds tags derived from your metadata, so you can filter and group that spend in Cost Reports, Virtual Tags, Budgets, and Cost Alerts. Because the telemetry comes from your gateway, this surfaces attribution the provider bill never itemized, even when many applications share one API key. LiteLLM Enrichment attributes costs for the model providers your proxy routes to and that you have connected as Vantage cost integrations: OpenAI, Anthropic, AWS Bedrock, Google Cloud (Vertex AI Gemini and Marketplace Claude), Azure, SpaceXAI, and Baseten. Vantage automatically routes each usage record to the matching provider’s costs.
Unlike Custom LLM Enrichment, Cloudflare AI Gateway Enrichment, and AWS Bedrock LLM Enrichment, which read logs you deliver to an S3 bucket you own, LiteLLM Enrichment runs a Vantage-managed collector next to your proxy that pushes usage to Vantage. There is no customer S3 bucket to create and no AWS role to grant.
Enrichment is metadata-only. Vantage reads the provider, model, token counts, and the tags you attach to each request. It does not collect your LiteLLM prompt or completion content or credentials, and no such content is written to a Vantage-owned artifact. This data is not used to train any models.

How It Works

The Vantage callback runs inside your LiteLLM process and emits one usage record per request to a local Vantage collector over a Unix socket. The collector compacts those records into batches, uploads them to a Vantage-managed bucket using short-lived presigned URLs, and Vantage joins that usage to your provider costs during each provider’s cost ingestion. See How LLM Enrichment works for the shared indexing, join, split, and tag pipeline.

How Cost Rows Are Split

Vantage splits each matching provider cost row into enriched rows, allocated proportionally by token share. Splits are additive, so the enriched rows always reconcile to the original total: usage your logs do not cover stays on a leftover row, and cost rows with no matching usage pass through unsplit. Enabling enrichment never changes your totals.

How Cost Splitting Works

See the allocation formula, a worked example, and how leftover rows keep totals reconciled.

Data Freshness and Backfill

Enrichment runs as part of each provider’s existing cost ingestion, so it follows that provider’s refresh cadence. Recent days are reprocessed within a rolling three-day window so late-arriving batches are picked up. See the provider data refresh documentation for per-provider timing. A billing period is enriched whenever it is processed while an active collector exists, for as long as the matching batches remain available to Vantage. Enrichment begins from when your collector starts uploading usage; it does not backfill spend from before the collector was running. Late-arriving batches for recent days are picked up automatically within the rolling three-day window. Re-enrichment reads the already-normalized cost data, so it does not require a full cost re-import.

Prerequisites

Before you begin, make sure:
  • An active cost integration exists for at least one supported provider: OpenAI, Anthropic, AWS, Google Cloud, Azure, SpaceXAI, or Baseten.
  • Your Azure deployments in LiteLLM, if any, use the azure/ (Azure OpenAI) or azure_ai/ (Azure AI Foundry) model prefix. Vantage routes usage by the provider LiteLLM reports, so a deployment configured with the openai/ prefix is reported as OpenAI even when its base URL points at Azure, and its usage will not match your Azure costs.
  • You run a LiteLLM proxy where you can install the Vantage callback and run the collector as a sidecar or adjacent container. The collector needs a persistent filesystem for its spool and outbound HTTPS access to Vantage Core and the presigned Amazon S3 URLs it returns.
  • You have a Vantage API token for the collector, in addition to the installation token you create in Step 1.
  • You have a Vantage Organization Owner or Integration Owner role. See Role-Based Access Control.

Set Up LiteLLM Enrichment

Setup has three steps: create the source and copy the installation token, install the callback and deploy the collector, and configure which cost integrations to enrich.

Step 1: Create the Source and Copy the Installation Token

In Vantage, go to the Integrations page. Under LLM Enrichment, add LiteLLM, then select Connect Source. Vantage provisions a collector integration and shows a Collector Setup screen. On that screen, copy the installation token from Step 1: Copy the Installation Token. It looks something like srvc_data_intgrtn_9f8e7d6c5b4a3f21. This token authorizes your collector to upload usage to Vantage. You paste it into your collector configuration in the next step.
The installation token identifies this collector integration to Vantage. Treat it as a secret and supply it to the collector through your platform’s secret management rather than committing it to a task definition or config file in plaintext.

Step 2: Install the Callback and Deploy the Collector

LiteLLM Enrichment has two runtime pieces that work together:
  • Vantage callback: A Python package loaded inside your LiteLLM proxy process that records each request’s usage.
  • Vantage collector: A sidecar container that receives events from the callback over a Unix socket, batches them, and uploads them to Vantage.
1

Install and enable the Vantage callback in LiteLLM

Install the vantage-litellm-callback package into the same Python environment as your LiteLLM proxy (the callback requires Python 3.10 or later), then register it so LiteLLM loads it.Install the callback package, pinning the version that matches your collector image (the current release is 0.0.7):
Official LiteLLM images may ship a virtual environment without pip on PATH. If pip is not found, bootstrap it first, then install into the same interpreter:
Register the callback in your LiteLLM configuration:
config.yaml
Installing the package makes vantage_callback importable, so the callback path above resolves on its own. Some LiteLLM images only load callbacks from Python files next to config.yaml. If LiteLLM does not pick up the callback after install, place a shim at /app/vantage_callback.py (the same directory as config.yaml) that imports callback_instance.
The callback is fail-open: collector outages, dropped acknowledgements, or queue pressure never block or fail a provider request. Usage events are buffered in a bounded in-process queue and delivered by a background worker, so enabling it does not add a hard dependency to your request path.
2

Deploy the Vantage collector sidecar

Run the Vantage collector as a sidecar or adjacent container in the same environment as your LiteLLM proxy, using the published image on Quay, pinned to a full version:
A release tag vX.Y.Z publishes the image tags X.Y.Z, X.Y, X, latest, and an immutable commit-SHA tag. Pin a full version tag (X.Y.Z) rather than latest; the current release is 0.0.7.
Run the collector image and the vantage-litellm-callback package on the same version, and upgrade them together. The callback-to-collector socket protocol and the batch schema are versioned in lockstep, so a mismatched pair can drop or reject usage.
Give the collector a persistent local spool, a Unix socket shared with LiteLLM, and outbound network access to Vantage. Configure it through environment variables (or your platform’s equivalent, such as a Helm value or a command-line flag):A minimal collector configuration looks like this:
Set VANTAGE_COLLECTOR_SOCKET_PATH to the same socket path on the LiteLLM container so the callback and collector connect.
The same installation token can be reused across multiple LiteLLM deployments and across multiple collector replicas behind one proxy. Each collector instance still needs its own stable VANTAGE_COLLECTOR_ID and its own persistent spool so Vantage can keep each batch stream distinct.
3

Restart LiteLLM

Restart or redeploy your LiteLLM proxy so the callback loads and usage records begin flowing to the collector. The collector buffers usage locally and uploads a batch once it fills, reaches its size limit, or crosses a UTC-day boundary (by default, about every 10 minutes, offset slightly per collector so replicas do not upload at the same moment), so the first upload can lag the restart.

Container Deployment Constraints

The release collector image is built FROM scratch, has no shell, runs as UID/GID 65532, and declares volumes for its spool and socket. These properties shape how you run it on any container platform; the notes below call out what that means, with ECS Fargate as a worked example.
Do not add shell-based health checks to the collector container. Because the image has no shell, CMD-SHELL, wget, and curl checks will not work. Instead, probe the collector’s HTTP health endpoint, GET /healthz, on its metrics address (:9090 by default, configurable with VANTAGE_METRICS_ADDRESS).
  • Health and metrics endpoints: The collector serves GET /healthz for readiness and Prometheus GET /metrics on its metrics address (:9090 by default). Use these for orchestrator health checks and monitoring. These endpoints are unauthenticated, so do not expose the metrics port outside your private network.
  • Memory: Request at least 256 MiB for the collector container and set a 512 MiB limit unless your own load testing supports a different value. The collector sizes its internal buffers from the container’s memory limit.
  • Per-task isolation: Each task or pod needs its own socket and spool volumes; do not share a spool across tasks. Use a stable VANTAGE_COLLECTOR_ID per instance, such as the task hostname.
  • Shared socket: Mount the same socket volume into both the LiteLLM and collector containers, and set VANTAGE_COLLECTOR_SOCKET_PATH (LiteLLM) and VANTAGE_SOCKET (collector) to that socket.
  • Secrets: Supply VANTAGE_API_TOKEN and VANTAGE_INTEGRATION_TOKEN through a secret store (for example, AWS Secrets Manager), not as plaintext values in a task definition or Compose file.
  • Volume ownership (ECS Fargate): Because the collector runs as UID 65532, and Fargate’s ephemeral volumes are root-owned, the collector may be unable to write its socket and spool by default. Override the collector task’s user to "0", or wrap the image with an entrypoint that chowns the mounts before dropping privileges.
For production, follow LiteLLM’s production deployment guidance: run two or more stateless proxy replicas behind a load balancer, each with its own collector and persistent collector state, and route traffic only through the load balancer.
On hosts that run LiteLLM outside a container (for example, on a VM), you can run the collector as a systemd service instead of a sidecar: install the binary, create a non-login vantage user, and provide the same environment variables through a mode-0600 environment file. The socket, spool, tokens, and provider selection all work the same way.

Step 3: Configure Cost Integrations

Back on the Collector Setup screen in Vantage, under Configure Cost Integrations, select Configure Providers and choose which connected cost integrations should receive enrichment from this collector.
Unlike the other LLM Enrichment integrations, which scan your logs and detect providers for you, LiteLLM requires you to choose exactly one cost integration per provider. LiteLLM usage records do not include a provider account identifier, so Vantage cannot tell apart multiple integrations for the same provider and will not enrich a provider that has more than one integration selected.
Vantage enriches costs for the selected integrations on each provider’s next data refresh. Provider cost integrations you connect later are not enriched automatically; return to the source’s Manage Providers screen to enable them.
On the Collector Setup screen, the Data Received indicator shows Waiting until your first usage arrives, then Received, usually within several minutes of your first traffic after the restart. If it shows Unavailable, see the Troubleshooting section.

Manage LiteLLM Enrichment Sources

Manage your collector from the LiteLLM integration page. The Connected Collectors table lists each collector with its Installation Token, Cost Providers, Status (for example, Pending before an import starts, Importing while one runs, Stable once all imports succeed, Warning if only some imports fail, Error if an import fails, or Paused if the collector is stopped), and creation date. Each row has an Edit button and an ellipses (…) menu with View import history and, for collectors that are not paused, Stop. The Collector Setup (Edit) screen also shows a Data Received indicator (Waiting, Received, or Unavailable) so you can confirm the collector is delivering usage independently of whether enrichment has run yet.

Choose Which Integrations Are Enriched

Select Edit on the collector, then Manage Providers, to change which connected cost integrations receive enrichment. Remember that only one cost integration per provider can be enriched. Disabling an integration here stops going-forward enrichment for it but leaves its existing enriched history in place.

View Import History

In the sources table, click the ellipses (…) next to a row and select View import history. This opens the collector’s Import History, where the Enrichment Runs table shows every run for this collector. The Enrichment Runs table lists one row per provider cost integration and billing period that ran enrichment, newest first, with columns for the Integration (and its account), Status, Tokens Kept, Log Lines Kept, Bill Match, Billing Period, and Last Enriched At. Lifecycle changes appear as their own rows: Added when enrichment is first enabled for an integration, and Paused or Resumed when you stop, disable, resume, or re-enable it. The three percentages look similar but measure different things at different stages, so a low number in one column means something very different from a low number in another: A low percentage does not change your totals. Usage Vantage cannot attribute stays on a leftover row, the unallocated remainder of the original cost row, so the enriched rows always add back up to what the provider billed. Select the Tokens Kept percentage on any run to open Token Metrics, which breaks that run down by disposition across the two stages and lets you view sample records for anything that did not make it through.

LLM Enrichment Import History

Full reference for the Import History screen: every column and status, how to read the three percentages, the Log Scan and Join dispositions with a fix for each, sample record fields, and a troubleshooting table.

Stop a Collector

To stop a collector, click the ellipses (…) next to the row and select Stop, then confirm in the Stop enrichment source dialog. This is a soft deactivation, not a hard delete: Vantage stops enriching new cost data for that collector, but your existing enriched history is left unchanged, and the collector stays visible in the list so you can bring it back at any time.
Stopping a collector does not re-import or roll back already-enriched costs. Going-forward enrichment simply pauses until you resume it. Stopping in Vantage does not stop the collector process itself. To stop sending usage, also stop the collector container.

Resume a Collector

To resume a stopped collector, select Resume in its row. Resuming restarts going-forward enrichment for the integrations that were active when you stopped it (integrations you had already disabled individually stay disabled) and reopens the Edit screen so you can adjust the selection.

View Enriched Costs on Cost Reports

Once enrichment runs, a single provider cost line is split into multiple rows, each carrying enrichment tags. You can filter and group by these tags anywhere tags are supported: Cost Reports, Virtual Tags, Budgets, and Cost Alerts. Because enrichment splits (allocates) your provider costs, enrichment tags behave like Vantage’s cost allocation tags: you can build a Virtual Tag on them, but a cost can be allocated only once, so an enrichment tag can belong to only one allocation chain.

Enrichment Tag Reference

The callback builds each usage record’s tags from your LiteLLM request metadata. It merges the metadata.tags and metadata.spend_logs_metadata objects (when both set the same key, metadata.tags wins), and metadata.tags may also be a list of key:value strings. It also promotes these LiteLLM request attributes to tags automatically, under the same key names: organization_id, organization_alias, team_id, team_alias, project_id, project_alias, user_id, end_user_id, and key_alias. Tag keys must be strings, and values must be strings, booleans, integers, or floats (null, nested, and array values are dropped). Each record carries at most 32 tags, with keys up to 128 characters and string values up to 256 characters. Your own metadata keys appear as tags without a provider prefix. The fields Vantage derives and adds itself, such as the model identifier, carry a vntg:ai: prefix (for example, vntg:ai:model); this keeps them distinct from your keys and consistent with the AI tags Vantage applies to provider costs. Because a key like team_alias is not provider-namespaced, it lines up across OpenAI, Anthropic, and the other providers, so you can group your entire AI stack by one team_alias tag. Consistent key naming matters: team and Team are two different keys.
Enriched rows retain the provider’s existing tags and add the Vantage-managed vntg:ai:* tags (such as vntg:ai:model); split rows also carry the tags from their usage slice. When a slice tag and an existing tag use the same key, the enrichment value wins. The leftover row (usage not covered by the collector) keeps the provider’s existing tags and the applicable vntg:ai:* tags but carries none of the slice tags. Request identifiers are never turned into tags. Vantage may also apply a tag denylist to drop noisy or sensitive keys before they become tags.
Enrichment tags behave like any other provider tag in the console:
  • To group: open the Group By menu, select Tag, and choose the tag key, for example vntg:ai:model, or one of your own keys such as team.
  • To filter: open the Filters menu, click New Rule, select Tag, choose the Tag Key, then pick an operator and one or more values.
The tag keys appear in the Tag Key dropdown once enriched costs exist. If you use a handful of keys often, mark them as preferred tags so they sort to the top of these menus.

Troubleshooting

The collector has not uploaded usage yet. Confirm that:
  • Your application is sending traffic through the LiteLLM proxy, and LiteLLM was restarted after you added the callback.
  • The callback is registered in litellm_settings.callbacks as vantage_callback.callback_instance and is importable in the LiteLLM process (installed with pip, or mounted so /app is on the import path).
  • The collector container is running, has outbound network access to Vantage, and shares the Unix socket with LiteLLM (VANTAGE_COLLECTOR_SOCKET_PATH matches on both). If the collector exits right after it starts, confirm VANTAGE_CORE_URL begins with https://.
  • VANTAGE_INTEGRATION_TOKEN matches the installation token from Step 1, and VANTAGE_API_TOKEN is a valid Vantage API token with integration permissions.
Vantage could not confirm the collector’s uploads. This usually means the collector cannot reach Vantage or its tokens are rejected. Confirm outbound HTTPS access, that VANTAGE_CORE_URL uses https://, and that VANTAGE_API_TOKEN and VANTAGE_INTEGRATION_TOKEN are set correctly. If it persists, contact Vantage Support.
Pending means Vantage has not run an enrichment import for this source yet. Enrichment runs during each provider’s regular cost ingestion, so a newly connected source shows Pending until the next data refresh of a provider it enriches. See the provider data refresh documentation for per-provider timing. If the status is still Pending after that refresh, check the following:
  • The Data Received indicator on the collector’s Collector Setup screen. Waiting means the collector has not uploaded usage yet: finish Step 2 and see the Waiting entry above.
  • If Data Received shows Received, confirm that at least one cost integration is selected under Configure Providers (Step 3), with exactly one integration per provider.
  • Confirm the provider has exactly one cost integration selected on the Manage Providers screen. A provider with more than one integration selected is not enriched because LiteLLM usage does not carry an account identifier.
  • Confirm the provider is one of the supported providers and that its cost integration is active.
  • For Azure, confirm your LiteLLM deployments use the azure/ or azure_ai/ model prefix. Usage from an openai/ deployment that points at Azure is reported as OpenAI. See Prerequisites.
  • Allow a full data refresh for that provider to complete after connecting. Enrichment applies on the next refresh.
Enrichment can only split costs by dimensions present in your usage. Costs pass through without allocation splits when any of the following is true:
  • There is no active collector for the account, or usage has not been ingested for that billing period.
  • The cost row’s model or token type could not be matched to a logged request for that provider, date, and token kind.
  • Records were skipped during ingestion because a required field was missing, the provider was blank or unsupported, the model was blank or could not be matched to your provider cost data, the request was not successful, or no usage value was a positive integer.
Confirm the request went to a supported provider, carries a matching model, and has at least one positive usage count.
Enrichment tags appear in the console as regular provider tags. To confirm what was enriched:
  • Check coverage in Import History: Open the source’s Import History and review Tokens Kept (the share of your logged tokens Vantage attached to a cost row) and Bill Match (the share of eligible provider cost rows that received logged usage) for each integration and billing period.
  • Check a Cost Report: Filter to the provider, then group by one of your own tag keys. Costs from requests that carried that key show its values. The leftover row (the portion of the bill your usage did not cover) and any usage sent without that key appear as Not tagged with the key name.
Totals should not change. Splits are additive and always sum to the original cost row to the cent, even when multiple sources are connected. If a total appears different, this is not expected behavior; contact Vantage Support.

Use Cases

Each row below shows example dimensions to attach to your LiteLLM requests and what that attribution enables in Vantage. Attach custom dimensions through LiteLLM’s metadata.tags or metadata.spend_logs_metadata; LiteLLM’s organization, team, project, user, end-user, and key attributes are captured automatically as organization_id, organization_alias, team_id, team_alias, project_id, project_alias, user_id, end_user_id, and key_alias. The key names are examples; you choose your own. All rows assume the collector is running and delivering usage.

Frequently Asked Questions

OpenAI, Anthropic, AWS Bedrock, Google Cloud (Vertex AI Gemini and Marketplace Claude), Azure, SpaceXAI, and Baseten. Vantage routes each usage record to the matching provider’s costs based on the provider your proxy called. Azure support covers direct Azure integrations; Azure CSP billing accounts are not supported.
Both attribute LLM spend from gateway telemetry. With Custom LLM Enrichment, you emit records in the Token Cost Allocation Specification and deliver them to an S3 bucket you own, and Vantage reads that bucket. With LiteLLM Enrichment, Vantage runs a managed collector next to your LiteLLM proxy that pushes usage to Vantage, so there is no bucket to create and no AWS role to grant. Use LiteLLM Enrichment when your traffic runs through a LiteLLM proxy.
No. You add the Vantage callback to your LiteLLM proxy configuration and run the collector sidecar; your application keeps calling LiteLLM as it does today. To attribute costs by your own dimensions, attach metadata to your LiteLLM requests as you normally would.
Yes. The same installation token can be reused across multiple LiteLLM deployments and across multiple collector replicas behind one proxy. Each collector instance needs its own stable VANTAGE_COLLECTOR_ID and its own persistent spool so Vantage can keep each batch stream distinct.
You need the Organization Owner or Integration Owner role to connect the source. Viewing the enriched costs does not require a special role; the resulting tags follow the same access rules as other cost data. See Role-Based Access Control.
No. Splits are additive and always sum to the original cost row. Existing provider-level reports continue to show the same totals; enrichment only makes new dimensions available on the underlying rows.
No. Enrichment is metadata-only. The collector sends the provider, model, token counts, and the tags you attach. Prompt and completion content and credentials are never collected, stored, or written to a Vantage-owned artifact.