Skip to content

Guides

The guides are Jupyter notebooks that take WarpRec apart one piece at a time, on real public data. Each one downloads its dataset, prepares it, and runs every step twice: through the Python API, where you can see what each object holds, and through the configuration file the pipelines read. The notebooks are committed with their outputs, so they can be read on GitHub without running anything.

They live in the guides/ folder of the repository, one folder per guide, each with its notebook and the configuration files it runs.

Running them

git clone https://github.com/sisinflab/warprec.git
cd warprec
pip install warprec jupyter        # guide 17 needs "warprec[mcp]"
jupyter lab guides/

Open each notebook from its own folder: the configuration files refer to the data relative to it. Datasets are downloaded from their publishers into guides/data/ the first time a guide needs them, and everything the guides write goes to guides/runs/; neither is committed. Most guides use MovieLens-100K and run in a minute or two on a laptop CPU; the ones that run the training pipeline on a local Ray instance or start a server (10, 11, 16 and 17) take five to fifteen minutes.

The guides

# Guide What it covers
Data
1 Reading data Files WarpRec reads, the reader, the Dataset, ready-made splits and folds
2 Filtering All thirteen filters, and why their order matters
3 Splitting All eight strategies, alignment, seeds, validation holdout and folds
4 Side information and clusters Item features, user and item clusters, the stash
5 Context Contextual fields, Transactions, duplicates, contextual evaluation
6 Knowledge graphs reader.knowledge, coverage, the KnowledgeGraph, KaHFM
7 Multimodal features reader.multimodal, file layouts, coverage, VBPR
8 Dataloaders Every training and evaluation loader, negative sampling, seeding
Models and training
9 Model families The registry, the model contract, one model per family, the design pipeline, checkpoints
10 Hyperparameter search The train pipeline on Ray: search spaces, strategies, schedulers, cross-validation
11 Running experiments Run names, pause and resume, dashboards, the estimate and swarm pipelines
Evaluation
12 Evaluation The Evaluator, full and sampled, mask_seen, every metric family, significance
13 Cold start, debiasing and re-ranking Cold-start protocols, IPS and SNIPS estimators, MMR and Calibration
14 Evaluating existing models The eval pipeline on checkpoints, ProxyRecommender, saved recommendations
Extending and deploying
15 Custom models and metrics Your own model, parameter class and metric through custom_modules
16 Callbacks and custom pipelines Every callback hook, where it runs, and a pipeline written against the API
17 Serving Serving checkpoints on Ray Serve: REST, sequential and contextual requests, MCP

The guides build on each other in this order, but each one runs on its own. The Ray cluster configuration for Google Cloud sits next to them; Cluster Management explains it.