Serverless AI Inference: Deploy Custom AI Model on Baseten with Fast Inference
What You'll Learn
- How Baseten and Truss fit into a custom model deployment
- Which deployment steps are documented by Baseten
- How model APIs and dedicated GPU pricing differ
- What to test before exposing an inference endpoint
What Baseten is for
Baseten is an inference platform for open-source, fine-tuned and custom models. Its pricing page describes a Basic plan at $0 per month with pay-as-you-go pricing and lists dedicated deployments, model APIs, training, fast cold starts and support options. The exact service terms depend on the account and deployment type. Compare the model-platform tradeoffs in this deployment guide before choosing a route.
The platform is not a substitute for model evaluation. A deployment can be technically successful while the model gives poor answers, uses too much memory or fails under the application's real request pattern. Compare the integration with the wider AI model pricing guide before choosing a provider.
What Truss does
Baseten's model documentation says custom models are packaged with Truss, its open-source model packaging tool. The first-model guide describes a configuration-led flow. A developer selects a model, chooses a GPU and deploys it to receive an endpoint.
Truss gives the deployment a repeatable description of the model, runtime and serving behavior. Keep that configuration in version control. Record the model revision, Python dependencies, tokenizer files, environment variables and expected input schema so a later deployment can be compared with the tested one.
Deploy custom AI model on Baseten: the documented flow
| Stage | What to decide | Evidence to keep |
|---|---|---|
| Model selection | Weights, license, task and input format | Model identifier and revision |
| Truss configuration | Runtime, dependencies, resources and handler | Committed config and build log |
| GPU selection | Memory, throughput and cost target | Test results and instance choice |
| Deployment | Endpoint, access policy and scaling | Deployment ID and API settings |
| Production test | Quality, latency, errors and spend | Representative request report |
Baseten's documentation describes deployment from a config file. That does not remove the need to validate dependencies and input handling. A model that works in a notebook may fail when the serving container receives a different tensor shape, an unsupported data type or a long prompt. Use the coding-model evaluation guide as a separate test-design reference.
Model APIs versus dedicated deployments
Baseten presents two different paths. Model APIs provide access to pre-optimized models running on its inference stack. Dedicated deployments run a selected model on allocated compute. A custom model generally needs the dedicated deployment path, while a supported catalog model may be available through a model API.
Do not compare a per-token model API rate with a dedicated GPU price as if they were the same product. The first is tied to token usage. The second is tied to compute time and deployment behavior. Your application may also need storage, network transfer, retries, logging and review costs.
Current Baseten model API pricing
| Model API listed by Baseten | Input per 1M tokens | Cache input per 1M tokens | Output per 1M tokens |
|---|---|---|---|
| Kimi K2.7 Code | $0.95 | $0.16 | $4.00 |
| Kimi K2.6 | $0.95 | $0.16 | $4.00 |
| GPT OSS 120B | $0.10 | Not listed | $0.50 |
These rows come from the Baseten pricing page fetched on August 22, 2026. The page labels the table as price per 1M tokens. Treat the values as a current snapshot, not a permanent contract. Check the pricing page and account terms before a launch.
Dedicated GPU pricing snapshot
Baseten's pricing page says dedicated deployments are billed for the compute used, down to the minute. The fetched table listed T4 at $0.01052 per minute, L4 at $0.01414, A10G at $0.02012, A100 at $0.06667 and H100 at $0.10833. The page also listed a B200 at $0.16633 per minute.
A simple four-hour estimate for one H100 at the listed rate is 240 multiplied by $0.10833, or $25.9992 before applicable account charges. That is a calculation, not a quote. Replicas, idle time, scaling behavior and regional availability can change actual spend.
How autoscaling changes the bill
Autoscaling can add capacity when demand rises and remove it when demand falls. Baseten's pricing page describes its Pro plan as including unlimited autoscaling and priority compute access. Basic, Pro and Enterprise have different support, compute and control terms, so choose the plan after checking the current product page.
Autoscaling is not the same as zero cost. A service may scale up during traffic spikes, keep a warm replica for response time or spend time building a deployment. Monitor active replicas, compute minutes, request volume, errors and cold starts together.
Cold starts and inference latency
Baseten lists fast cold starts as a Basic feature, but no universal latency figure should be promised for every custom model. Startup time depends on image size, weights, runtime, GPU allocation, model initialization and traffic pattern. Measure p50 and p95 latency on the exact model and configuration you plan to deploy.
Keep a warm-up test separate from steady-state inference. Report the request size, output length, concurrency, region, model revision and GPU. This makes the result useful to another engineer instead of turning one local test into a broad performance claim.
Designing the prediction endpoint
Baseten's concepts documentation says custom model deployments use the predict API, which accepts and returns application data. Define the request and response schema before deployment. Validate required fields, reject unexpected sizes and return clear error messages.
Keep authentication outside the model prompt. Use short-lived credentials where possible, rotate secrets and avoid writing sensitive request bodies to logs. Add request IDs so an application error can be linked to the endpoint log without exposing private content.
Model memory and GPU selection
GPU choice starts with model memory, but it should not end there. Include weights, runtime overhead, tokenizer memory, activation memory, batching and the maximum input or output size. A model that fits during a single request may fail when several requests arrive together.
Start with the smallest instance that passes a representative load test. Then compare the cost of a larger GPU against the cost of queuing, lower throughput or more replicas. Keep the result tied to a specific model revision because memory behavior can change after an update.
Security checks for a custom model
Review the model license and the data path before deployment. Decide whether prompts, outputs, traces and logs contain personal or confidential information. Use a test dataset that does not contain live secrets, and document who can read the deployment logs.
Restrict tool access if the model is connected to external actions. A prediction endpoint that only returns text has a different risk profile from an agent that can call a database, shell command or payment system. The agentic AI controls guide explains why permissions should be narrow.
Observability and cost control
| Metric | Why it matters | Action when it changes |
|---|---|---|
| Requests and tokens | Shows usage and billing drivers | Set budget alerts |
| GPU minutes | Shows dedicated compute use | Review replicas and idle time |
| p50 and p95 latency | Separates normal from slow requests | Check batching and GPU size |
| Error and timeout rate | Shows reliability problems | Inspect input limits and runtime logs |
Log enough information to reproduce a failure without storing more user data than necessary. A monthly review should connect traffic, compute, errors and accepted outputs. The Kimi coding-model cost guide uses the same cost-per-use approach.
Baseten deployment checklist
Before launch, pin the model revision, test the Truss build, verify the input schema, select a GPU with headroom, configure authentication and record the rollback path. Run a small load test with realistic prompts. Confirm that the endpoint returns errors safely and that logs do not expose secrets.
After launch, monitor cost and quality together. Set a threshold for error rate, latency and daily spend. Re-test after changing the model, runtime, GPU, tokenizer or autoscaling settings. A deployment is a maintained service, not a one-time upload.
Deploy custom AI model on Baseten decision guide
Baseten is a practical option to test when you want a managed route for open-source, fine-tuned or custom model inference. Truss provides a configuration-led packaging path, while dedicated deployments and model APIs serve different cost and control needs. The current pricing page supports pay-as-you-go deployment and lists token and GPU rates that should be rechecked before launch.
Use a controlled pilot to decide. Record model quality, p50 and p95 latency, cold starts, GPU minutes, cache or token usage, errors and review time. Do not carry forward unsupported fixed-speed, zero-scaling or universal cost claims without a current test.
Frequently Asked Questions
SK Jabedul Haque
Building India's most trusted finance education platform — simplifying news, schemes and market trends so anyone can understand and invest confidently.
Read full bioNever miss an update
Get our clearest explainers on schemes, markets and money — read what matters, without the noise.
Explore more articles