Machine Learning Infrastructure

ML Platform Architecture

Ray for distributed compute, JupyterHub for notebooks, and MLflow for tracking — scheduled onto GPU nodes that Karpenter spins up and tears down on demand.

ML platform architecture diagram
01

Architecture Decisions

Scaling

Karpenter automatically scales nodes up or down depending on utilization, resulting in lower compute costs — important when GPU nodes are involved.

Node optimization

Taints/tolerations and node selectors make sure Ray and Jupyter pods land on the compute-heavy and GPU nodes rather than general-purpose ones.

IP allocation

Ray and JupyterHub were split into separate subnets, so if either scales past its allocated IPs the whole cluster doesn't get clogged up.

Service accounts

A Karpenter service account to spin up EC2 nodes, and an MLflow service account to interact with the MLflow database and the S3 artifact store.

ML platform Istio routing
02

Services I Used

Ray

Ray on the cluster

Ray is an open-source unified framework for scaling AI and Python applications like machine learning.

  • Model serving — a simple Python API that lets users compose complex inference pipelines and deploy them in real time using YAML.
  • Computation — Ray is superior to Spark for machine learning workloads.
  • Libraries — scalable libraries for common ML tasks: data preprocessing, distributed training, hyperparameter tuning, reinforcement learning, and model serving.
  • More info on Ray here.

JupyterHub

JupyterHub

JupyterHub gives users access to computational environments and resources without burdening them with installation and maintenance. Users work in their own workspaces on shared resources that administrators can manage efficiently.

  • Collaboration — multiple users can collaborate on the same notebook.
  • Customization — the deployed notebook image can be designed per role, whether for a data engineer (Spark) or a data scientist (Ray).

MLflow

MLflow

MLflow provides a suite of tools aimed at simplifying the ML workflow, tailored to assist practitioners throughout the stages of ML development and deployment.

  • Experiment management — easy to keep track of all experiments.
  • Model management — storing the model in a registry allows for reproducibility when deploying.
03

Consoles

Ray Cluster Dashboard

/ray-cluster-dashboard — web-based dashboard for monitoring and debugging Ray applications.

Ray Serve Dashboard

/ray-service-dashboard — monitoring and debugging for the Ray Serve deployment.

Ray Serve Endpoint

/ray-service-serve — API endpoint that turns input data into predictions.

Jupyter

/jupyter — web-based notebook for writing code in multiple languages.

The cluster underneath — EKS, Istio mTLS, the Prometheus/Grafana stack, and ArgoCD delivery — is documented on the Kubernetes page.