Serving Models¶
A model saved by the training pipeline can be served directly from its checkpoint with one command and one configuration file. Serving runs on Ray Serve:
- concurrent requests to a model are batched into one forward pass;
- every model runs in its own replicas, which can be scaled by hand or by an autoscaling policy;
- models can share a GPU, each asking for a fraction of it;
- the same configuration runs on a laptop, joins an existing Ray cluster, or is exported for
serve deployand KubeRay.
General, sequential, graph-based and context-aware models are all served from the checkpoint alone. The dataset, the experiment directory and the training configuration are not needed.
Install¶
pip install "warprec[serving]" # the REST API
pip install "warprec[mcp]" # the REST API and the MCP endpoint for LLM agents
Serving a model needs the same extras as training it: graph-based models such as LightGCN or NGCF also need graph, for example pip install "warprec[serving,graph]". A missing one shows up as a ModuleNotFoundError in the replica's log, and the server stops with Deploying application warprec failed.
From Training to Serving¶
1. Save the model. Set save_model in the model's meta section of the training configuration:
The best model of the run is written to {local_experiment_path}/{dataset_name}/serialized/{Model}_{name}.pth. Besides the weights, the file carries everything serving needs from the training data: the items each user has seen, each user's recent history for a sequential model, and the context vocabulary for a context-aware one.
2. Describe what to serve in a serving configuration (see Serving Configuration for every key):
server:
port: 8000
api_key: change-me
endpoints:
- name: sasrec
checkpoint: experiments/ml-1m/serialized/SASRec_brave-otter.pth
item_metadata:
path: data/ml-1m/movies.dat
sep: "::"
header: false
encoding: latin-1
3. Serve it.
The command blocks until it is stopped. Ctrl-C at a terminal and SIGTERM from a process manager or a container both shut it down cleanly.
API¶
Every route but /healthz requires the X-API-Key header when an API key is configured.
| Method | Path | Purpose |
|---|---|---|
| GET | /healthz |
Liveness check. |
| GET | /v1/models |
The cards of all served models. |
| GET | /v1/models/{name} |
One model's card: how to ask it, what it knows, how it was trained. |
| GET | /v1/models/{name}/context |
The context a context-aware model accepts, with an example. |
| GET | /v1/models/{name}/items?q= |
Items whose name contains q, with their attributes. |
| POST | /v1/models/{name}/items/lookup |
Items by id or name. |
| POST | /v1/models/{name}/popular |
The items with the most training interactions. |
| POST | /v1/models/{name}/recommend |
Top-k recommendations. |
| POST | /v1/models/{name}/score |
Scores of given candidate items, for re-ranking. |
| * | /mcp |
MCP tools, resources and prompts, when server.mcp is true. |
List the models.
[{"name": "sasrec", "model": "SASRec", "kind": "sequential", "n_users": 6040, "n_items": 3706,
"needs_user": false, "warprec_version": "1.16.0", "params": {"embedding_size": 64, "...": "..."},
"context": null}]
Recommend for a known user. Ids are the dataset's own, as numbers or strings.
curl -X POST localhost:8000/v1/models/sasrec/recommend -H "X-API-Key: change-me" \
-H "Content-Type: application/json" -d '{"user_id": 1, "k": 3}'
{"model": "sasrec", "fallback": false,
"items": [{"item_id": 2858, "score": 7.91, "name": "American Beauty (1999)"},
{"item_id": 1196, "score": 7.64, "name": "Star Wars: Episode V - The Empire Strikes Back (1980)"},
{"item_id": 1210, "score": 7.48, "name": "Star Wars: Episode VI - Return of the Jedi (1983)"}]}
Recommend for a session. A sequential model also answers from a history of items, oldest first, without a known user. With an item catalogue, items may be given by their exact name.
curl -X POST localhost:8000/v1/models/sasrec/recommend -H "X-API-Key: change-me" \
-H "Content-Type: application/json" \
-d '{"history": ["Toy Story (1995)", "Aladdin (1992)"], "k": 5, "exclude": [2355]}'
Score candidates that another system retrieved. Scores come back in request order, and nothing is masked.
curl -X POST localhost:8000/v1/models/sasrec/score -H "X-API-Key: change-me" \
-H "Content-Type: application/json" -d '{"user_id": 1, "items": [1, 260, 2858]}'
{"model": "sasrec", "scores": [{"item_id": 1, "score": 5.12, "name": "Toy Story (1995)"},
{"item_id": 260, "score": 7.55, "name": "Star Wars: Episode IV - A New Hope (1977)"},
{"item_id": 2858, "score": 7.91, "name": "American Beauty (1999)"}]}
Errors carry a detail message and a status: 404 for an unknown model or user, 422 for a request the model cannot answer, 401 for a missing or wrong API key.
What a Model Says About Itself¶
Every endpoint describes itself, so a client or an agent needs nothing but the server to use it.
The model card, GET /v1/models/{name}, holds:
how_to_ask: what a request must and may contain, in sentences.example_request: a request the endpoint answers, to start from.catalogue: whether items have names, and the attributes they carry with their most common values.training: what the model was trained on and how it scored.- dataset, date, and numbers of users, items and interactions;
- evaluation strategy and test metrics, such as
nDCG@10.
contextandexample_context: for a context-aware model.- The configured
descriptionanditem_noun, the model class, its hyperparameters and the WarpRec version.
The catalogue. With item_metadata, items have names and attributes, such as genres. These are returned with every item.
- Names are matched ignoring case.
- A name that is not in the catalogue gets a 422 suggesting the closest ones ("did you mean 'Toy Story (1995)'?").
- Search answers whether a model knows an item:
curl -s -H "X-API-Key: change-me" "localhost:8000/v1/models/sasrec/items?q=toy%20story"
curl -s -H "X-API-Key: change-me" "localhost:8000/v1/models/sasrec/items?q=Kung%20Fu%20Panda" # no matches: not in its training data
curl -s -X POST localhost:8000/v1/models/sasrec/items/lookup -H "X-API-Key: change-me" \
-H "Content-Type: application/json" -d '{"items": ["heat (1995)", 260]}'
The context. GET /v1/models/{name}/context answers "what context can I give you?", with three things per field:
- its accepted values, most frequent in training first, with their counts;
- the training range, for a numeric field;
- the configured description.
It also gives an example context that the model accepts.
Checkpoints saved before WarpRec recorded training facts still serve. Their card has training: null, and their context values come without counts.
Popular Items, Filters and Explanations¶
Popular items. POST /v1/models/{name}/popular lists the items with the most training interactions. It is the same for everyone, and suits a first visit:
curl -s -X POST localhost:8000/v1/models/sasrec/popular -H "X-API-Key: change-me" \
-H "Content-Type: application/json" -d '{"k": 5, "filter": {"genres": "Comedy"}}'
Filters. recommend and popular accept filter, such as {"genres": "Comedy"}.
- Only items with those attributes are returned, ranked as the model ranks them.
- Values are compared ignoring case.
- A list matches any of its values, for example
{"genres": ["Action", "Crime"]}. - Several attributes must all match.
- An unknown attribute or value gets a 422 that lists or suggests the accepted ones.
Explanations. explain: true on recommend adds because to each item: up to two of the request's own items (the user's training items, or the session) that training users most often consumed together with it, each with co_occurrences, the number of training users who consumed both. For user 1 above, the first item becomes:
{"item_id": 2858, "score": 7.91, "name": "American Beauty (1999)",
"because": [{"item_id": 608, "name": "Fargo (1996)", "co_occurrences": 1840},
{"item_id": 2762, "name": "Sixth Sense, The (1999)", "co_occurrences": 1787}]}
Warning
because is evidence from the training data, not the model's reasoning: WarpRec models do not expose why they rank an item. Present it as "people who liked X also liked this", not as the reason the model chose it.
How Requests Are Answered¶
- Ids. Requests and responses use the dataset's own user and item ids. They are matched as strings, so
1and"1"are the same user. - Seen items. With
mask_seen(the default), items the user interacted with in training are left out, as are the items of a request'shistoryandexclude. When masking leaves fewer thankitems, fewer are returned: a masked item is never used to pad the list. - Sequential models.
- A
user_idalone uses the user's training history. - A
historyalone is an anonymous session. - Both together score the given history.
- Caser, FOSSIL and STAN mix a user embedding into the sequence, so they need a known
user_idwith a history (needs_user: truein the model listing). - A
historysent to a non-sequential model is refused.
- A
- Unknown users. A user absent from the training data gets
404, or, withunknown_user: popular, the items most interacted with in training andfallback: truein the response. - Context-aware models (AFM, DCN, DCNv2, DeepFM, FM, NFM, Wide&Deep, xDeepFM) need the situation of every request.
contextis a dictionary by field, or a list in the field order the model listing shows.- The listing also shows each field's type and the values it knows.
- A categorical field takes one value seen in training, a multi-valued field takes a list of them, and a numeric field takes a number.
- A value never seen in training is refused, listing the known ones, rather than silently scored as unknown.
- Seen-item masking stays per user, not per context, as when recommendations are written to file.
curl -X POST localhost:8000/v1/models/fm/recommend -H "X-API-Key: change-me" \
-H "Content-Type: application/json" \
-d '{"user_id": 12, "k": 5, "context": {"daytime": "morning", "weather": "sunny", "isweekend": "workday", "homework": "home"}}'
MCP¶
With server.mcp: true and the mcp extra installed, the same server speaks the Model Context Protocol under /mcp. The API key protects it as well.
Instructions. On connecting, a client receives instructions it passes to its model before any tool is called. They state what the server does, list each endpoint with its configured description, and say which tool answers which question.
Tools:
| Tool | Answers |
|---|---|
list_models |
Which models are there, what each serves and needs? |
describe_model |
How do I ask this model? What does it know, and how was it trained? |
describe_context |
What context can I provide, and which values are accepted? |
search_items |
Is this model trained on "Kung Fu Panda"? Which "Toy Story" films does it know? |
get_items |
What are these items, and what attributes do they have? |
popular_items |
What is popular, possibly among comedies only? |
recommend |
What should this user, or this session, try next, possibly filtered and explained? |
score_items |
Which of these candidates would the user like most? |
Resources: warprec://models (the list) and warprec://models/{name} (a model card), for clients that browse resources instead of calling tools.
Prompts:
- recommend_for_me guides a conversation that ends in explained, personal recommendations.
- explore_catalogue guides an exploration of what a model knows.
A client configuration looks like this:
{"mcpServers": {"warprec": {"url": "http://localhost:8000/mcp/", "headers": {"X-API-Key": "change-me"}}}}
Scaling and GPUs¶
Each endpoint is its own Ray Serve deployment, configured independently:
- batching: concurrent requests are grouped into one forward pass of up to
max_batch_sizerequests, waiting at mostbatch_wait_timeout_sfor a batch to fill. A replica accepts as many requests at once as a batch holds, unlessdeployment.max_ongoing_requestssays otherwise. - deployment.num_replicas fixes the number of replicas, while deployment.autoscaling_config lets Ray Serve add and remove them with the load.
- deployment.ray_actor_options sets the resources of each replica. With
device: cudaa replica asks for one GPU;num_gpus: 0.25places four replicas on one GPU instead. - server.ray_address: auto joins the Ray cluster the machine belongs to, rather than starting a local one.
Deploying¶
warprec.serve runs the same way everywhere Ray runs:
- one machine;
- a Docker container;
- a Kubernetes Deployment;
- a Ray cluster on virtual machines;
- a KubeRay RayService.
For a cluster, --export writes the application as a standard Ray Serve config file:
Deploying Served Models walks through every option, with the image, the manifests and the trade-offs.
Security¶
- Checkpoints are pickles. Loading one can run code, so only serve files from a trusted source.
- Set an API key whenever the server is reachable beyond the local machine. The
WARPREC_API_KEYenvironment variable keeps it out of configuration files.
Limitations¶
- Checkpoints saved before this version serve without seen-item masking or popularity fallback. Graph-based and context-aware models among them must be saved again.
- Checkpoints are read from the local file system.