Machine Learning Infrastructure
Ray for distributed compute, JupyterHub for notebooks, and MLflow for tracking — scheduled onto GPU nodes that Karpenter spins up and tears down on demand.
Karpenter automatically scales nodes up or down depending on utilization, resulting in lower compute costs — important when GPU nodes are involved.
Taints/tolerations and node selectors make sure Ray and Jupyter pods land on the compute-heavy and GPU nodes rather than general-purpose ones.
Ray and JupyterHub were split into separate subnets, so if either scales past its allocated IPs the whole cluster doesn't get clogged up.
A Karpenter service account to spin up EC2 nodes, and an MLflow service account to interact with the MLflow database and the S3 artifact store.
Ray is an open-source unified framework for scaling AI and Python applications like machine learning.

JupyterHub gives users access to computational environments and resources without burdening them with installation and maintenance. Users work in their own workspaces on shared resources that administrators can manage efficiently.

MLflow provides a suite of tools aimed at simplifying the ML workflow, tailored to assist practitioners throughout the stages of ML development and deployment.
/ray-cluster-dashboard — web-based dashboard for monitoring and debugging Ray applications.
/ray-service-dashboard — monitoring and debugging for the Ray Serve deployment.
/ray-service-serve — API endpoint that turns input data into predictions.
/jupyter — web-based notebook for writing code in multiple languages.
The cluster underneath — EKS, Istio mTLS, the Prometheus/Grafana stack, and ArgoCD delivery — is documented on the Kubernetes page.