Python API¶
Signatures for the public objects. Guides: Python for calling a server, Docker or UV / Pip for running one.
Client¶
Import from tserve.client. Needs the client extra.
Client ¶
Call a TServe inference server from Python.
The client connects to a local TServe server. Use it as a context manager so the HTTP session is closed.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
url
|
str
|
Server origin, for example |
required |
timeout
|
float
|
Request timeout in seconds. Ignored if |
60.0
|
transport
|
BaseTransport or None
|
Optional prebuilt transport. When omitted, uses HTTP. |
None
|
Examples:
>>> from tserve.client import Client
>>> with Client("http://127.0.0.1:8000") as client:
... result = client.predict(
... past={"timestamp": ["2024-01-01", "2024-01-02"], "sales": [120, 135]},
... time="timestamp",
... target=["sales"],
... fh=3,
... model="naive",
... )
See Also
Python client Install, connect, predict, and use native table types. Install and serve Start the local server this client calls. HTTP API Send the same predict fields as JSON instead of Arrow.
Construct a Client.
predict ¶
predict(
*,
past: Any,
fh: int,
time: str | None = None,
target: str | list[str] | None = None,
model: str = "naive",
future: Any = None,
static: Any = None,
quantiles: list[float] | None = None,
) -> PredictResponse
Send a prediction and restore tables to the type of past.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
past
|
any
|
Past observations (dict, pandas, polars, …). Return type
of |
required |
fh
|
int
|
Prediction steps ahead ( |
required |
time
|
str
|
Time-index column. When omitted, the first column of
|
None
|
target
|
str or list of str
|
Columns to forecast. When omitted, every other |
None
|
model
|
str
|
Loaded model id (see |
``"naive"``
|
future
|
any
|
Known future values of time-varying covariates. Must include the time column and cover the forecast horizon. |
None
|
static
|
any
|
One row of values that stay constant over time. Broadcast
across history and the forecast. Does not require |
None
|
quantiles
|
list of float
|
Quantile alphas, e.g. |
None
|
Returns:
| Type | Description |
|---|---|
PredictResponse
|
|
Raises:
| Type | Description |
|---|---|
ValidationError
|
If the request shape or columns are invalid. |
RuntimeError
|
If the server returns HTTP >= 400. |
Notes
This method sends Arrow tables to POST /predict/bytes.
JSON clients use POST /predict.
See Also
Data specification Table formats, column roles, inference, and response fields. Models catalog Available model ids and their dependencies. Errors Local validation, transport, and server failures.
health ¶
health() -> HealthResult
Return GET /health (liveness).
Returns:
| Type | Description |
|---|---|
HealthResult
|
Currently |
See Also
HTTP status routes Route behavior and response examples.
models ¶
models() -> ModelsResult
Return GET /models — loaded ids, not the registry catalog.
Returns:
| Type | Description |
|---|---|
ModelsResult
|
Each row has |
See Also
Models catalog Model ids the server can load. HTTP status routes Loaded-model response semantics.
stats ¶
stats() -> StatsResult
Return GET /stats — uptime, memory, per-model latency.
Returns:
| Type | Description |
|---|---|
StatsResult
|
Process and per-loaded-id metrics. |
See Also
HTTP status routes Stats response fields and example.
close ¶
Close the HTTP session.
Called automatically when a with Client(...) block exits.
See Also
Python client Context-manager usage.
Server¶
Import from tserve.server. Needs the server extra and a family extra for the models you load.
Server ¶
Server(
model: list[str | Path | tuple[str, Any]] | None = None,
models_dir: str | Path | None = None,
*,
host: str = "127.0.0.1",
port: int = 8000,
log_level: str = "info",
)
Run a TServe inference HTTP server.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model
|
list of str or (str, object)
|
Extra registry ids to load, |
None
|
models_dir
|
str or Path
|
Directory of saved sktime |
None
|
host
|
str
|
Bind address. |
``"127.0.0.1"``
|
port
|
int
|
Bind port. |
8000
|
log_level
|
str
|
Uvicorn log level. |
``"info"``
|
Attributes:
| Name | Type | Description |
|---|---|---|
url |
str
|
|
app |
FastAPI
|
The ASGI app (dashboard, JSON, OpenAPI). |
Raises:
| Type | Description |
|---|---|
ValueError
|
Unknown registry id, empty craft spec, non-zip path, or duplicate id. |
TypeError
|
|
ImportError
|
Executor extra is not installed. |
Examples:
>>> from tserve.server import Server
>>> Server(host="127.0.0.1", port=8000).run()
>>> Server(
... model=[
... "chronos_bolt",
... ("drift", 'NaiveForecaster(strategy="drift")'),
... ]
... )
Notes
naive is always loaded as a test baseline. Extra ids in model
load alongside it for real forecasts.
See Also
Install and serve
Installation, startup options, and server URLs.
Docker
Image defaults, tags, and container arguments.
Models catalog
Registry ids and required family extras.
Live objects
In-process (id, estimator) pairs.
Craft specs
(id, spec) pairs and CLI id=spec.
Dashboard
Browser interface served by app at GET /.
Construct a Server.
url
property
¶
url: str
Return http://{host}:{port}.
See Also
Install and serve Bind addresses, ports, and published routes.
run ¶
Serve self.app with uvicorn.
Blocks until the process exits.
See Also
From source
Run the server from Python or tserve.
HTTP API
Routes exposed by the running app.
Dashboard
Browser interface at GET /.
Predict request and response¶
PredictRequest is both the JSON body of POST /predict and the keyword signature of Client.predict. Table formats and column rules are in the data specification.
PredictRequest ¶
Predict input. Same fields on JSON POST /predict and Client.predict.
Tables may be a column dict (name → list), a row matrix
{"columns": [...], "data": [[...], ...]}, pandas, polars,
pyarrow, or Narwhals.
model is a loaded id (see GET /models), not an executor
name such as "sktime".
Attributes:
| Name | Type | Description |
|---|---|---|
past |
any
|
Past observations. After coercion must include |
time |
(str or None, optional)
|
Time-index column. When omitted, the first column of |
target |
str or list of str or None, optional
|
Target columns. A string becomes a one-element list. When
omitted, remaining |
fh |
int
|
Prediction steps ahead (must be |
model |
str, default ``"naive"``
|
Loaded model id. |
future |
(any, optional)
|
Future values of time-varying covariates, over the forecast
horizon. Non-target columns shared with |
static |
(any, optional)
|
One-row static features, broadcast over time. |
quantiles |
list of float, optional
|
Quantile alphas, e.g. |
Notes
past is a table, and time names one of its columns.
See Also
Data specification
Table formats, column roles, and inference rules.
HTTP predict route
JSON request behavior and status codes.
Python client
Arrow transport through Client.predict.
PredictResponse ¶
Predict output.
JSON POST /predict returns tables as column dicts.
Client.predict restores the type of the caller's past.
Attributes:
| Name | Type | Description |
|---|---|---|
predictions |
any
|
Point-forecast table, including the time column. |
model |
str
|
Model id that produced the forecast. |
request_id |
str
|
Server-assigned UUID for this call. |
quantiles |
(any, optional)
|
Quantile table when requested; column names are
|
See Also
Data specification
JSON and native-table response behavior.
HTTP predict route
JSON response and failure statuses.
Python client
Access the result returned by Client.predict.
Status payloads¶
Returned by client.health(), client.models(), and client.stats(), and by the matching HTTP routes.
HealthResult ¶
Payload for GET /health.
Attributes:
| Name | Type | Description |
|---|---|---|
status |
str
|
Currently always |
error |
HealthError or None
|
Unused on the live route. |
See Also
HTTP status routes Health semantics and response shape.
HealthError ¶
Structured error nested under HealthResult.error.
This is a response body, not a raised exception. The current
GET /health handler always returns status="ok" with
error omitted.
Attributes:
| Name | Type | Description |
|---|---|---|
code |
str
|
Machine-readable error code (example: |
message |
str
|
Human-readable explanation. |
See Also
Errors HTTP statuses and Python exceptions. HTTP status routes Current health-route behavior.
ModelsResult ¶
Payload for GET /models.
Attributes:
| Name | Type | Description |
|---|---|---|
models |
list of ModelInfo
|
Currently loaded models. Empty if nothing was loaded. |
See Also
Models catalog
Available registry ids; this result contains only loaded ids.
HTTP status routes
GET /models response behavior.
ModelInfo ¶
One loaded model (a row of GET /models).
Attributes:
| Name | Type | Description |
|---|---|---|
id |
str
|
Id used in predict |
executor |
{sktime, pytorch - forecasting, custom}
|
Plugin that loaded the artifact. |
source |
{object, registry, directory, craft}
|
Registry id, saved |
See Also
Models catalog Registry ids the server can load. HTTP status routes Loaded-model listing and source values.
StatsResult ¶
Payload for GET /stats.
Attributes:
| Name | Type | Description |
|---|---|---|
uptime_s |
float
|
Seconds since process start. |
memory |
MemoryStats
|
RSS and GPU probes; fields may be |
models |
dict of str to ModelStats
|
Per loaded model id load, warmup, and latency. |
See Also
HTTP status routes Stats semantics and response example.
MemoryStats ¶
Process memory snapshot nested under StatsResult.memory.
Attributes:
| Name | Type | Description |
|---|---|---|
cpu_rss_mb |
float or None
|
Resident set size in MiB, or |
gpu_mb |
float or None
|
CUDA memory allocated across devices in MiB, or |
See Also
HTTP status routes
Full GET /stats response.
ModelStats ¶
Per-model load, warmup, and traffic nested under StatsResult.models.
Keys of StatsResult.models are model ids, not executor names.
Attributes:
| Name | Type | Description |
|---|---|---|
executor |
str or None
|
Executor plugin name recorded at load time. |
load_s |
float or None
|
Wall time of |
warmup_s |
float or None
|
Wall time of |
requests |
RequestCounts
|
Success/failure counts from |
latency_s |
LatencySummary
|
Predict-call wall times, including failed calls. |
See Also
HTTP status routes
Full GET /stats response.
RequestCounts ¶
Per-model request counters nested under ModelStats.requests.
Attributes:
| Name | Type | Description |
|---|---|---|
total |
int
|
|
ok |
int
|
Predict calls that returned without raising. |
failed |
int
|
Predict calls that raised (still counted in latency). |
See Also
HTTP status routes
Full GET /stats response.
LatencySummary ¶
Per-model predict latency nested under ModelStats.latency_s.
Attributes:
| Name | Type | Description |
|---|---|---|
count |
int
|
Number of |
total |
float
|
Sum of wall times in seconds. |
mean |
float or None
|
|
fastest |
float or None
|
Minimum recorded duration, or |
slowest |
float or None
|
Maximum recorded duration, or |
See Also
HTTP status routes
Full GET /stats response.