Skip to main content

Workload Sizing Guides

Use this catalog to choose the Lakemeter form that matches the consumption you need to estimate. Each linked guide is the canonical source for that workload's sizing inputs, calculation behavior, and Excel export.

For current Databricks product capabilities and availability, refer to the official Databricks documentation. These guides intentionally focus on how Lakemeter models cost.

Adding a workload — select type, configure parameters, and save Choose a workload type, enter its sizing assumptions, and save it to the estimate.

Compute and SQL

What you need to sizeLakemeter guideMain sizing inputs
Scheduled or triggered processingLakeflow JobsCompute shape, workers, runs, runtime
Interactive notebook computeAll-Purpose ComputeCompute shape, workers, active hours
Declarative data pipelinesLakeflow Spark Declarative PipelinesCompute mode, edition, workers, usage
Managed connector pipelines through the APILakeflow Connect APIPipeline usage, edition, optional gateway
SQL warehousesDatabricks SQLWarehouse type, size, clusters, hours

AI, ML, and data services

What you need to sizeLakemeter guideMain sizing inputs
Model inference endpointsModel ServingEndpoint type, scale-out concurrency, hours
Vector indexing and searchAI SearchEndpoint mode, vector capacity, storage, reranker requests, hours
Databricks-hosted foundation modelsFoundation Models — DatabricksModel, rate type, token volume or hours
Proprietary foundation modelsFoundation Models — ProprietaryProvider, model, geography, context, token volume
Transactional database capacityLakebaseCompute range, usage, nodes, storage protection
Databricks-hosted applicationsDatabricks AppsApp size, app count, active hours
Document parsingAI ParseEstimation mode, document complexity, page volume
Structured field extractionAI ExtractDocument type, monthly input volume
Document classificationAI ClassifyDocument type, monthly document volume
Gateway inference tables and usage trackingUnity AI GatewayComponents enabled, input method, request or payload volume
Agent evaluation service usageAgent EvaluationFeatures enabled, evaluation token volume, synthetic questions
Serverless GPU model trainingAI RuntimeAccelerator, monthly runtime
Image generationShutterstock ImageAIMonthly image volume

Ingestion, storage, and platform

What you need to sizeLakemeter guideMain sizing inputs
Direct or OpenTelemetry ingestionZerobus IngestMode, monthly ingested GB
Databricks-managed default storageDatabricks Default StorageStored data, Tier 1 operations, Tier 2 operations
Security and mission-critical packagesPlatform Add-onsAdd-on selection, cloud, tier, Product Spend at List

Shared sizing principles

Name the assumption

Use a workload name that identifies the scenario being modeled, not only the product. For example, Nightly customer ingestion is easier to review than Jobs workload.

Model expected usage

Lakemeter may ask for run frequency and duration, active hours, capacity, storage, tokens, pages, or images. Use representative usage rather than maximum technical limits unless the estimate is intentionally modeling a peak case.

Use the options shown in the app

Available clouds, regions, tiers, sizes, SKUs, and models can change. The current Lakemeter controls and pricing data determine what can be estimated. Check the Databricks pricing page and your commercial terms before using an estimate for a final purchasing decision.

Review the breakdown

Expand each workload after saving it. Confirm the usage quantity, billing unit, selected SKU, rate, and any separate VM or storage components before exporting.

  • SKU Explorer — inspect SKU list rates and model simple volume-based costs
  • FMAPI Tokens — compare proprietary model token-rate combinations
  • Calculation Reference — understand the calculation structure shared across workloads
  • Exporting to Excel — understand how sizing assumptions appear in the exported workbook