Skip to content

Python API

Signatures for the public objects. Guides: Python for calling a server, Docker or UV / Pip for running one.

Client

Import from tserve.client. Needs the client extra.

Client

Client(
    url: str,
    *,
    timeout: float = 60.0,
    transport: BaseTransport | None = None,
)

Call a TServe inference server from Python.

The client connects to a local TServe server. Use it as a context manager so the HTTP session is closed.

Parameters:

Name Type Description Default
url str

Server origin, for example "http://127.0.0.1:8000". Ignored if transport is given.

required
timeout float

Request timeout in seconds. Ignored if transport is given.

60.0
transport BaseTransport or None

Optional prebuilt transport. When omitted, uses HTTP.

None

Examples:

>>> from tserve.client import Client
>>> with Client("http://127.0.0.1:8000") as client:
...     result = client.predict(
...         past={"timestamp": ["2024-01-01", "2024-01-02"], "sales": [120, 135]},
...         time="timestamp",
...         target=["sales"],
...         fh=3,
...         model="naive",
...     )
See Also

Python client Install, connect, predict, and use native table types. Install and serve Start the local server this client calls. HTTP API Send the same predict fields as JSON instead of Arrow.

Construct a Client.

predict

predict(
    *,
    past: Any,
    fh: int,
    time: str | None = None,
    target: str | list[str] | None = None,
    model: str = "naive",
    future: Any = None,
    static: Any = None,
    quantiles: list[float] | None = None,
) -> PredictResponse

Send a prediction and restore tables to the type of past.

Parameters:

Name Type Description Default
past any

Past observations (dict, pandas, polars, …). Return type of predictions matches this.

required
fh int

Prediction steps ahead (> 0).

required
time str

Time-index column. When omitted, the first column of past is used.

None
target str or list of str

Columns to forecast. When omitted, every other past column is a target, except columns also present in future.

None
model str

Loaded model id (see GET /models).

``"naive"``
future any

Known future values of time-varying covariates. Must include the time column and cover the forecast horizon.

None
static any

One row of values that stay constant over time. Broadcast across history and the forecast. Does not require future.

None
quantiles list of float

Quantile alphas, e.g. [0.1, 0.5, 0.9].

None

Returns:

Type Description
PredictResponse

predictions, optional quantiles, model, and request_id.

Raises:

Type Description
ValidationError

If the request shape or columns are invalid.

RuntimeError

If the server returns HTTP >= 400.

Notes

This method sends Arrow tables to POST /predict/bytes. JSON clients use POST /predict.

See Also

Data specification Table formats, column roles, inference, and response fields. Models catalog Available model ids and their dependencies. Errors Local validation, transport, and server failures.

health

health() -> HealthResult

Return GET /health (liveness).

Returns:

Type Description
HealthResult

Currently status='ok'.

See Also

HTTP status routes Route behavior and response examples.

models

models() -> ModelsResult

Return GET /models — loaded ids, not the registry catalog.

Returns:

Type Description
ModelsResult

Each row has id, executor, and source.

See Also

Models catalog Model ids the server can load. HTTP status routes Loaded-model response semantics.

stats

stats() -> StatsResult

Return GET /stats — uptime, memory, per-model latency.

Returns:

Type Description
StatsResult

Process and per-loaded-id metrics.

See Also

HTTP status routes Stats response fields and example.

close

close() -> None

Close the HTTP session.

Called automatically when a with Client(...) block exits.

See Also

Python client Context-manager usage.

Server

Import from tserve.server. Needs the server extra and a family extra for the models you load.

Server

Server(
    model: list[str | Path | tuple[str, Any]] | None = None,
    models_dir: str | Path | None = None,
    *,
    host: str = "127.0.0.1",
    port: int = 8000,
    log_level: str = "info",
)

Run a TServe inference HTTP server.

Parameters:

Name Type Description Default
model list of str or (str, object)

Extra registry ids to load, (id, estimator) pairs, or (id, craft spec) string pairs. naive is always loaded as a test baseline. Default [] loads only naive. When models_dir is set, matching .zip stems already in this list are loaded from disk.

None
models_dir str or Path

Directory of saved sktime .zip files. Not loaded wholesale.

None
host str

Bind address.

``"127.0.0.1"``
port int

Bind port.

8000
log_level str

Uvicorn log level.

``"info"``

Attributes:

Name Type Description
url str

http://{host}:{port}.

app FastAPI

The ASGI app (dashboard, JSON, OpenAPI).

Raises:

Type Description
ValueError

Unknown registry id, empty craft spec, non-zip path, or duplicate id.

TypeError

(id, object) whose object is neither a craft spec string nor a sktime BaseForecaster.

ImportError

Executor extra is not installed.

Examples:

>>> from tserve.server import Server
>>> Server(host="127.0.0.1", port=8000).run()
>>> Server(
...     model=[
...         "chronos_bolt",
...         ("drift", 'NaiveForecaster(strategy="drift")'),
...     ]
... )
Notes

naive is always loaded as a test baseline. Extra ids in model load alongside it for real forecasts.

See Also

Install and serve Installation, startup options, and server URLs. Docker Image defaults, tags, and container arguments. Models catalog Registry ids and required family extras. Live objects In-process (id, estimator) pairs. Craft specs (id, spec) pairs and CLI id=spec. Dashboard Browser interface served by app at GET /.

Construct a Server.

url property

url: str

Return http://{host}:{port}.

See Also

Install and serve Bind addresses, ports, and published routes.

run

run() -> None

Serve self.app with uvicorn.

Blocks until the process exits.

See Also

From source Run the server from Python or tserve. HTTP API Routes exposed by the running app. Dashboard Browser interface at GET /.

Predict request and response

PredictRequest is both the JSON body of POST /predict and the keyword signature of Client.predict. Table formats and column rules are in the data specification.

PredictRequest

Predict input. Same fields on JSON POST /predict and Client.predict.

Tables may be a column dict (name → list), a row matrix {"columns": [...], "data": [[...], ...]}, pandas, polars, pyarrow, or Narwhals.

model is a loaded id (see GET /models), not an executor name such as "sktime".

Attributes:

Name Type Description
past any

Past observations. After coercion must include time and every target.

time (str or None, optional)

Time-index column. When omitted, the first column of past.

target str or list of str or None, optional

Target columns. A string becomes a one-element list. When omitted, remaining past columns (minus future columns).

fh int

Prediction steps ahead (must be > 0).

model str, default ``"naive"``

Loaded model id.

future (any, optional)

Future values of time-varying covariates, over the forecast horizon. Non-target columns shared with past are passed to the estimator as exogenous features.

static (any, optional)

One-row static features, broadcast over time.

quantiles list of float, optional

Quantile alphas, e.g. [0.1, 0.5, 0.9].

Notes

past is a table, and time names one of its columns.

See Also

Data specification Table formats, column roles, and inference rules. HTTP predict route JSON request behavior and status codes. Python client Arrow transport through Client.predict.

PredictResponse

Predict output.

JSON POST /predict returns tables as column dicts. Client.predict restores the type of the caller's past.

Attributes:

Name Type Description
predictions any

Point-forecast table, including the time column.

model str

Model id that produced the forecast.

request_id str

Server-assigned UUID for this call.

quantiles (any, optional)

Quantile table when requested; column names are {variable}_{alpha}.

See Also

Data specification JSON and native-table response behavior. HTTP predict route JSON response and failure statuses. Python client Access the result returned by Client.predict.

Status payloads

Returned by client.health(), client.models(), and client.stats(), and by the matching HTTP routes.

HealthResult

Payload for GET /health.

Attributes:

Name Type Description
status str

Currently always "ok".

error HealthError or None

Unused on the live route.

See Also

HTTP status routes Health semantics and response shape.

HealthError

Structured error nested under HealthResult.error.

This is a response body, not a raised exception. The current GET /health handler always returns status="ok" with error omitted.

Attributes:

Name Type Description
code str

Machine-readable error code (example: MODEL_NOT_LOADED).

message str

Human-readable explanation.

See Also

Errors HTTP statuses and Python exceptions. HTTP status routes Current health-route behavior.

ModelsResult

Payload for GET /models.

Attributes:

Name Type Description
models list of ModelInfo

Currently loaded models. Empty if nothing was loaded.

See Also

Models catalog Available registry ids; this result contains only loaded ids. HTTP status routes GET /models response behavior.

ModelInfo

One loaded model (a row of GET /models).

Attributes:

Name Type Description
id str

Id used in predict model.

executor {sktime, pytorch - forecasting, custom}

Plugin that loaded the artifact.

source {object, registry, directory, craft}

Registry id, saved .zip, in-process estimator, or a user-supplied sktime craft spec.

See Also

Models catalog Registry ids the server can load. HTTP status routes Loaded-model listing and source values.

StatsResult

Payload for GET /stats.

Attributes:

Name Type Description
uptime_s float

Seconds since process start.

memory MemoryStats

RSS and GPU probes; fields may be None.

models dict of str to ModelStats

Per loaded model id load, warmup, and latency.

See Also

HTTP status routes Stats semantics and response example.

MemoryStats

Process memory snapshot nested under StatsResult.memory.

Attributes:

Name Type Description
cpu_rss_mb float or None

Resident set size in MiB, or None if probing failed.

gpu_mb float or None

CUDA memory allocated across devices in MiB, or None if torch/CUDA is unavailable.

See Also

HTTP status routes Full GET /stats response.

ModelStats

Per-model load, warmup, and traffic nested under StatsResult.models.

Keys of StatsResult.models are model ids, not executor names.

Attributes:

Name Type Description
executor str or None

Executor plugin name recorded at load time.

load_s float or None

Wall time of Executor.load.

warmup_s float or None

Wall time of Executor.warmup.

requests RequestCounts

Success/failure counts from Scheduler.run.

latency_s LatencySummary

Predict-call wall times, including failed calls.

See Also

HTTP status routes Full GET /stats response.

RequestCounts

Per-model request counters nested under ModelStats.requests.

Attributes:

Name Type Description
total int

ok + failed.

ok int

Predict calls that returned without raising.

failed int

Predict calls that raised (still counted in latency).

See Also

HTTP status routes Full GET /stats response.

LatencySummary

Per-model predict latency nested under ModelStats.latency_s.

Attributes:

Name Type Description
count int

Number of Scheduler.run calls recorded for this model.

total float

Sum of wall times in seconds.

mean float or None

total / count, or None when count is 0.

fastest float or None

Minimum recorded duration, or None when count is 0.

slowest float or None

Maximum recorded duration, or None when count is 0.

See Also

HTTP status routes Full GET /stats response.