Upload Markdown and plain-text files. Training reads them and writes the question–answer pairs it learns from; you start one run, not a preparation step.
Your documents,
trained into your own model.
Scalty reads what your team already wrote, writes training examples from it, fine-tunes an open-weight model and measures it against the base — then serves it behind an endpoint you own.
The loop
From a folder of documents to a model that answers like your team.
LoRA on a license-clean open-weight base. Settings are derived from your data or set by hand; time and price are estimated before the run starts.
Held-out questions from your own documents, answered by the base and by your model, side by side. The model card reports what was measured.
Training carries voice and terminology. For prices, rules and stock, attach a memory collection — answers come back with a citation into the source.
Answer in the playground, call an OpenAI-compatible endpoint with a key, or download the weights and run them on your own machine.
Every figure on the billing page comes from the usage ledger. $5 of credit every month; no plans, no seats.
Point the OpenAI SDK at your endpoint.
Streaming, usage and citations come back in the shapes you already handle. Keys are scoped to the organisation; every deployment has revisions you can roll back.
- chat/completions · predictions
- API keys per organisation, metered requests
- export-manifest: take the weights with you
from openai import OpenAI client = OpenAI( base_url="https://studio.scalty.com/api/v1", api_key="sk_live_…", ) r = client.chat.completions.create( model="support-assistant-prod", messages=[{"role": "user", "content": "Hi"}], )
Metered. Read straight off the ledger.
No plans, no seats. A run shows its cost while it works and its billed minutes when it ends; the balance page says exactly what was spent. The rates shown are today's rate book.