Skip to the content.

Deployment

The search service is a stateless FastAPI app (necs.api.app:app) that loads four artifacts at startup and serves /search.

Artifacts

Env var Default Produced by
NECS_RETRIEVER artifacts/bi_encoder train_retriever.sh
NECS_RERANKER artifacts/cross_encoder train_reranker.sh
NECS_INDEX artifacts/product_index scripts/build_index.py
NECS_CATALOGUE artifacts/catalogue.json scripts/build_index.py / export

catalogue.json is a flat {product_id: product_text} map used to fetch the passage shown to the reranker and returned to the client. If artifacts are missing the app still boots and /health reports model_loaded: false, so Kubernetes/compose readiness probes behave correctly.

Local (bare metal)

pip install -r requirements.txt
make serve            # uvicorn necs.api.app:app --host 0.0.0.0 --port 8000

Docker

# Build
docker build -t necs-search:latest .

# Run with trained artifacts mounted read-only
docker run --rm -p 8000:8000 \
    -v "$(pwd)/artifacts:/app/artifacts:ro" \
    necs-search:latest

Or with compose:

docker compose up --build

The image is CPU-only (faiss-cpu, CPU torch wheels). For GPU serving, base the image on pytorch/pytorch:*-cuda* and install faiss-gpu.

API

The JSON below is an illustrative response shape. Its catalogue size, latency, identifier, score, and text are synthetic examples, not measurements from a published deployment.

GET /health

{ "status": "ok", "model_loaded": true, "catalogue_size": 482105 }

POST /search

Request:

{ "query": "wireless gaming mouse", "top_k": 5 }

Response:

{
  "query": "wireless gaming mouse",
  "took_ms": 41.7,
  "results": [
    {
      "product_id": "B0XXXXXXX",
      "product_text": "logitech g pro x superlight wireless gaming mouse ...",
      "score": 0.94,
      "label": "E",
      "label_name": "Exact",
      "retrieval_rank": 2
    }
  ]
}

Try it:

curl -s localhost:8000/search \
  -H 'content-type: application/json' \
  -d '{"query":"wireless gaming mouse","top_k":5}' | jq

Interactive docs are served at http://localhost:8000/docs (Swagger UI).

Hosting a live demo

The service is a standard containerized web app, so any Docker host works. The one hard requirement is that trained artifacts must be available to the container (baked into the image or on a mounted disk) — otherwise /search returns 503 and /health reports model_loaded: false.

A public live endpoint is intentionally not committed here because it requires trained artifacts and hosted infrastructure. The configurations are starting points, not a verified one-click deployment; validate artifact loading, memory, latency, authentication, abuse controls, and costs before exposing the service publicly.

Documentation site

Project documentation is published with GitHub Pages from the docs/ folder. Once enabled (Settings → Pages → Source: main / /docs), it is served at https://madhvansh.github.io/Neural-E-Commerce-Search/.

Scaling notes