Skip to Content

Serverless AI Inference: Deploy Custom AI Model on Baseten with Fast Inference

How to Package Open-Source LLMs with Truss, Scale GPU Compute from Zero, and Serve Low-Latency APIs
2026-08-17 19:16:46 Updated 2026-08-22 05:58:59.693904 — min read 59 views
Serverless AI Inference: Deploy Custom AI Model on Baseten with Fast Inference
“Deploy custom AI model on Baseten with a Truss configuration, a selected GPU and a managed prediction endpoint. Baseten's current documentation covers open-source and custom model deployment, autoscaling and model APIs. This guide explains the documented flow, cost structure and production checks without promising a fixed latency.

What You'll Learn

  • How Baseten and Truss fit into a custom model deployment
  • Which deployment steps are documented by Baseten
  • How model APIs and dedicated GPU pricing differ
  • What to test before exposing an inference endpoint

What Baseten is for

Baseten is an inference platform for open-source, fine-tuned and custom models. Its pricing page describes a Basic plan at $0 per month with pay-as-you-go pricing and lists dedicated deployments, model APIs, training, fast cold starts and support options. The exact service terms depend on the account and deployment type. Compare the model-platform tradeoffs in this deployment guide before choosing a route.

The platform is not a substitute for model evaluation. A deployment can be technically successful while the model gives poor answers, uses too much memory or fails under the application's real request pattern. Compare the integration with the wider AI model pricing guide before choosing a provider.

What Truss does

Baseten's model documentation says custom models are packaged with Truss, its open-source model packaging tool. The first-model guide describes a configuration-led flow. A developer selects a model, chooses a GPU and deploys it to receive an endpoint.

Truss gives the deployment a repeatable description of the model, runtime and serving behavior. Keep that configuration in version control. Record the model revision, Python dependencies, tokenizer files, environment variables and expected input schema so a later deployment can be compared with the tested one.

Deploy custom AI model on Baseten: the documented flow

StageWhat to decideEvidence to keep
Model selectionWeights, license, task and input formatModel identifier and revision
Truss configurationRuntime, dependencies, resources and handlerCommitted config and build log
GPU selectionMemory, throughput and cost targetTest results and instance choice
DeploymentEndpoint, access policy and scalingDeployment ID and API settings
Production testQuality, latency, errors and spendRepresentative request report

Baseten's documentation describes deployment from a config file. That does not remove the need to validate dependencies and input handling. A model that works in a notebook may fail when the serving container receives a different tensor shape, an unsupported data type or a long prompt. Use the coding-model evaluation guide as a separate test-design reference.

Model APIs versus dedicated deployments

Baseten presents two different paths. Model APIs provide access to pre-optimized models running on its inference stack. Dedicated deployments run a selected model on allocated compute. A custom model generally needs the dedicated deployment path, while a supported catalog model may be available through a model API.

Do not compare a per-token model API rate with a dedicated GPU price as if they were the same product. The first is tied to token usage. The second is tied to compute time and deployment behavior. Your application may also need storage, network transfer, retries, logging and review costs.

Current Baseten model API pricing

Model API listed by BasetenInput per 1M tokensCache input per 1M tokensOutput per 1M tokens
Kimi K2.7 Code$0.95$0.16$4.00
Kimi K2.6$0.95$0.16$4.00
GPT OSS 120B$0.10Not listed$0.50

These rows come from the Baseten pricing page fetched on August 22, 2026. The page labels the table as price per 1M tokens. Treat the values as a current snapshot, not a permanent contract. Check the pricing page and account terms before a launch.

Dedicated GPU pricing snapshot

Baseten's pricing page says dedicated deployments are billed for the compute used, down to the minute. The fetched table listed T4 at $0.01052 per minute, L4 at $0.01414, A10G at $0.02012, A100 at $0.06667 and H100 at $0.10833. The page also listed a B200 at $0.16633 per minute.

A simple four-hour estimate for one H100 at the listed rate is 240 multiplied by $0.10833, or $25.9992 before applicable account charges. That is a calculation, not a quote. Replicas, idle time, scaling behavior and regional availability can change actual spend.

How autoscaling changes the bill

Autoscaling can add capacity when demand rises and remove it when demand falls. Baseten's pricing page describes its Pro plan as including unlimited autoscaling and priority compute access. Basic, Pro and Enterprise have different support, compute and control terms, so choose the plan after checking the current product page.

Autoscaling is not the same as zero cost. A service may scale up during traffic spikes, keep a warm replica for response time or spend time building a deployment. Monitor active replicas, compute minutes, request volume, errors and cold starts together.

Cold starts and inference latency

Baseten lists fast cold starts as a Basic feature, but no universal latency figure should be promised for every custom model. Startup time depends on image size, weights, runtime, GPU allocation, model initialization and traffic pattern. Measure p50 and p95 latency on the exact model and configuration you plan to deploy.

Keep a warm-up test separate from steady-state inference. Report the request size, output length, concurrency, region, model revision and GPU. This makes the result useful to another engineer instead of turning one local test into a broad performance claim.

Designing the prediction endpoint

Baseten's concepts documentation says custom model deployments use the predict API, which accepts and returns application data. Define the request and response schema before deployment. Validate required fields, reject unexpected sizes and return clear error messages.

Keep authentication outside the model prompt. Use short-lived credentials where possible, rotate secrets and avoid writing sensitive request bodies to logs. Add request IDs so an application error can be linked to the endpoint log without exposing private content.

Model memory and GPU selection

GPU choice starts with model memory, but it should not end there. Include weights, runtime overhead, tokenizer memory, activation memory, batching and the maximum input or output size. A model that fits during a single request may fail when several requests arrive together.

Start with the smallest instance that passes a representative load test. Then compare the cost of a larger GPU against the cost of queuing, lower throughput or more replicas. Keep the result tied to a specific model revision because memory behavior can change after an update.

Security checks for a custom model

Review the model license and the data path before deployment. Decide whether prompts, outputs, traces and logs contain personal or confidential information. Use a test dataset that does not contain live secrets, and document who can read the deployment logs.

Restrict tool access if the model is connected to external actions. A prediction endpoint that only returns text has a different risk profile from an agent that can call a database, shell command or payment system. The agentic AI controls guide explains why permissions should be narrow.

Observability and cost control

MetricWhy it mattersAction when it changes
Requests and tokensShows usage and billing driversSet budget alerts
GPU minutesShows dedicated compute useReview replicas and idle time
p50 and p95 latencySeparates normal from slow requestsCheck batching and GPU size
Error and timeout rateShows reliability problemsInspect input limits and runtime logs

Log enough information to reproduce a failure without storing more user data than necessary. A monthly review should connect traffic, compute, errors and accepted outputs. The Kimi coding-model cost guide uses the same cost-per-use approach.

Baseten deployment checklist

Before launch, pin the model revision, test the Truss build, verify the input schema, select a GPU with headroom, configure authentication and record the rollback path. Run a small load test with realistic prompts. Confirm that the endpoint returns errors safely and that logs do not expose secrets.

After launch, monitor cost and quality together. Set a threshold for error rate, latency and daily spend. Re-test after changing the model, runtime, GPU, tokenizer or autoscaling settings. A deployment is a maintained service, not a one-time upload.

Deploy custom AI model on Baseten decision guide

Baseten is a practical option to test when you want a managed route for open-source, fine-tuned or custom model inference. Truss provides a configuration-led packaging path, while dedicated deployments and model APIs serve different cost and control needs. The current pricing page supports pay-as-you-go deployment and lists token and GPU rates that should be rechecked before launch.

Use a controlled pilot to decide. Record model quality, p50 and p95 latency, cold starts, GPU minutes, cache or token usage, errors and review time. Do not carry forward unsupported fixed-speed, zero-scaling or universal cost claims without a current test.

Frequently Asked Questions

Baseten's documentation describes Truss as its open-source model packaging tool for custom model deployments. A Truss configuration describes the model and serving setup so the deployment can be repeated and reviewed.
The documented flow is to package the model with Truss, choose a GPU and deploy it from a configuration. The resulting custom deployment exposes a prediction endpoint that the application can call.
Baseten's pricing page lists a Basic plan at $0 per month with pay-as-you-go pricing. Dedicated deployments are billed for compute used down to the minute, while model APIs use token-based pricing.
The Baseten pricing page fetched for this update listed an H100 at $0.10833 per minute. This is a pricing snapshot, not a quote, and actual spend can vary with replicas, idle time, region, taxes and account terms.
A model API provides access to a supported pre-optimized model and is generally billed by tokens. A dedicated deployment runs a selected model on allocated compute and is billed by compute time and deployment behavior.
No universal latency should be assumed for every custom model. Startup time and inference depend on model size, runtime, GPU, input, output, concurrency and traffic. Measure p50 and p95 on the exact configuration.
Test the model revision, input schema, dependencies, GPU memory, quality, p50 and p95 latency, cold starts, errors, authentication, logs and cost. Use a controlled dataset and keep a rollback path.
SK Jabedul Haque
Written by

SK Jabedul Haque

Founder & Chief Editor

Building India's most trusted finance education platform — simplifying news, schemes and market trends so anyone can understand and invest confidently.

Read full bio

Never miss an update

Get our clearest explainers on schemes, markets and money — read what matters, without the noise.

Explore more articles
In this article