Articles

Model serving and compatible APIs: one endpoint, several contracts

A serving engine adds a queue, a lifecycle and an API surface to a weights file.

Reading: 6 minAutomation & AI

Article cover: Model serving and compatible APIs: one endpoint, several contracts

A weights file cannot answer a request. Something has to hold the model in memory, decide how many requests to run at once, queue the rest, and speak HTTP — that something is the serving engine, and it is where the phrase “OpenAI-compatible” does its work. The habit worth building is this: treat compatibility as a claim about shape, verify it against the engine’s own documentation, and remember that the same engine usually offers more than one contract at the same time.

What serving adds to a model

A model file plus a server is a different kind of object from a model file:

  • residency — the model stays loaded between requests; in Ollama the keep_alive parameter controls how long it stays in memory after the last request.
  • a queue and a batching strategy — server slots (-np, --parallel) and continuous batching (-cb, --cont-batching, enabled by default), which is what turns several requests into one batch.
  • a lifecycle — which model is loaded, when it is loaded, and what happens to the first request that arrives before it is ready.
  • an HTTP surface — endpoints, authentication, streaming, error shapes and, optionally, metrics.

The surface a local engine actually exposes

The llama.cpp server is a useful concrete case because its documentation lists every route. Its defaults are deliberately conservative: --host defaults to 127.0.0.1 and --port to 8080.

A health endpoint that needs no key. GET /health is documented as “public (no API key check)” with /v1/health as an alias, and it reports 503 while the model is not ready — which is precisely what a load balancer or a start-up script wants to poll.

OpenAI-compatible routes. GET /v1/models (“OpenAI-compatible Model Info API”), POST /v1/completions, POST /v1/chat/completions, POST /v1/embeddings and POST /v1/responses are each documented as OpenAI-compatible APIs.

A second contract. POST /v1/messages is documented as an “Anthropic-compatible Messages API” — the same server, the same loaded model, a different request and response shape.

Native routes. Alongside them the server exposes its own endpoints: /completion, /tokenize, /detokenize, /apply-template, /infill, /props, /slots, /lora-adapters. They are richer than the compatible ones — and, being engine-specific, they are not portable.

Observability and security. GET /metrics is a “Prometheus compatible metrics exporter”, disabled until --metrics is passed; authentication is off until you pass --api-key (a comma-separated list) or --api-key-file; TLS is optional (--ssl-key-file, --ssl-cert-file).

What “compatible” does and does not mean

It means the route and the JSON shape match closely enough that a client written for the OpenAI API works unchanged. The engine’s own documentation marks the seam explicitly: some endpoints are “not OAI-compatible” — /completion and /embeddings are described that way, next to their OpenAI-compatible counterparts /v1/completions and /v1/embeddings, which exist precisely because the response formats differ.

What the shape does not cover: default values for parameters the caller omits, fields the local engine adds to a response, the exact body and status of an error, rate-limit headers (a local engine has none in the provider’s sense — see the API sheet in this batch), model naming and aliases, and whether a parameter is implemented at all. Two engines can both be “compatible” and disagree about every one of those.

The consequence of running one

A local server is an HTTP service that, by default, listens only on loopback and accepts any caller. Making it reachable from the network is a second decision, and it is separate from the first: --host 0.0.0.0 publishes the port, while --api-key is what decides who may use it. Publishing without a key exposes the machine’s accelerator, its loaded model and its logs to anyone who can reach the port. The /health endpoint stays public by design, which is reasonable because it discloses only readiness — and it is worth knowing when reading an engine’s threat surface.

Level and prerequisites. L2 — operational: the reader must be able to bring a model up behind an HTTP interface, name what the serving layer adds, and check a compatibility claim instead of trusting it. No server is installed or configured in this sheet, and no deployment is described. Prerequisites: the memory and placement sheets of this batch.

Where to go next

References

  • ggml-org — llama.cpp server — the default host and port, the full endpoint list with each route’s own description, the API-key options, TLS options and the metrics endpoint.
  • Ollama — API reference — a second local engine’s surface: POST /api/generate and POST /api/chat on localhost:11434, with streaming by default.
  • vLLM — documentation — a serving engine described as exposing an “OpenAI-compatible API server, plus Anthropic Messages API and gRPC support”.
Nodrius Field Kit — cover

Get the Nodrius Field Kit

Enter your email and we'll send you the complete Field Kit: eight technical resources in one download.

Your email address is sent to Nodrius and used to deliver the Field Kit to you by email. Marketing messages are separate and optional: the checkbox does not affect your download, and you are only added to our updates list if you tick it. Privacy Policy.

Privacy & cookies

This site does not use cookies, analytics or tracking, and it does not profile you. There is nothing technical to switch off, so Accept and Reject change nothing about how the site works: either choice lets you browse everything normally.

The only difference is what your browser remembers: your choice is stored on this device so this notice is not shown again. It contains no identifier, it is not shared with anyone, and the site sets no cookie for it.

What we do with your data
How your email address is used when you ask for a Resource, how long it is kept and which rights you have — in the Privacy Policy.
What the site stores in your browser
Nothing: no cookies, no local storage, no third-party content — in the Cookie Policy.